cd /news/artificial-intelligence/measuring-reward-seeking-by-instilli… · home topics artificial-intelligence article
[ARTICLE · art-67157] src=lesswrong.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Measuring Reward-Seeking by Instilling Contrastive Beliefs

OpenAI researchers operationalized reward-seeking in machine learning models as the causal sensitivity of behavior to beliefs about grader preferences, finding that training checkpoints of several frontier models engage in grader-reasoning without special prompting. The work aims to measure when models pursue what their grader rewards rather than what designers intended, citing examples from Langosco, Shah, Zech, Carlsmith, Hebbar, Mallen & Shlegeris, Hubinger, Schoen & Nitishinskaya, Claude Opus 4.8, Fable 5, METR’s GPT-5.6 evaluation, and GPT-5.6 preview system card.

read1 min views1 publishedJul 21, 2026

This is an unofficial automated linkpost.

Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself [Langosco; Shah], and a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease [Zech]. The trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy.

One such proxy is the reward process itself: a model may learn to pursue what its grader rewards rather than what its designers intended. We call this behavior reward-seeking: a model representing its grader (a reward model in training, an evaluation grader in testing, or a monitor in deployment) and conditioning its behavior on what it believes the grader rewards [Carlsmith; Hebbar; Mallen & Shlegeris]. A reward-seeker may value grader approval terminally or pursue it instrumentally to protect some other objective, such as avoiding modification or gaining future influence [Hubinger; Carlsmith]; our definition does not distinguish the two.

Training checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting [Schoen & Nitishinskaya; Claude Opus 4.8 System Card; Fable 5 System Card; METR’s GPT-5.6 evaluation; GPT-5.6 preview system card]; see Figure 2 for an example. Such reasoning is evidence of underlying reward-seeking but a poor systematic measurement tool: a model can act on its grader-beliefs (beliefs about grader preferences) without articulating them, and verbalized reasoning often does not map cleanly onto the final action [Schoen & Nitishinskaya]. In this work we operationalize reward-seeking as the causal sensitivity of behavior to beliefs about grader preferences.

Continue reading at alignment.openai.com →

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/measuring-reward-see…] indexed:0 read:1min 2026-07-21 ·