{"slug": "measuring-reward-seeking-by-instilling-contrastive-beliefs", "title": "Measuring Reward-Seeking by Instilling Contrastive Beliefs", "summary": "OpenAI researchers operationalized reward-seeking in machine learning models as the causal sensitivity of behavior to beliefs about grader preferences, finding that training checkpoints of several frontier models engage in grader-reasoning without special prompting. The work aims to measure when models pursue what their grader rewards rather than what designers intended, citing examples from Langosco, Shah, Zech, Carlsmith, Hebbar, Mallen & Shlegeris, Hubinger, Schoen & Nitishinskaya, Claude Opus 4.8, Fable 5, METR’s GPT-5.6 evaluation, and GPT-5.6 preview system card.", "body_md": "*This is an unofficial automated linkpost.*\n\nMachine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself [[Langosco](https://arxiv.org/abs/2105.14111); [Shah](https://arxiv.org/abs/2210.01790)], and a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease [[Zech](https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1002683)]. The trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy.\n\nOne such proxy is the reward process itself: a model may learn to pursue what its grader rewards rather than what its designers intended. We call this behavior reward-seeking: a model representing its grader (a reward model in training, an evaluation grader in testing, or a monitor in deployment) and conditioning its behavior on what it believes the grader rewards [[Carlsmith](https://arxiv.org/abs/2311.08379); [Hebbar](https://blog.redwoodresearch.org/p/how-training-gamers-might-function); [Mallen & Shlegeris](https://www.lesswrong.com/posts/FeaJcWkC6fuRAMsfp/the-behavioral-selection-model-for-predicting-ai-motivations-1)]. A reward-seeker may value grader approval terminally or pursue it instrumentally to protect some other objective, such as avoiding modification or gaining future influence [[Hubinger](https://arxiv.org/abs/1906.01820); [Carlsmith](https://arxiv.org/abs/2311.08379)]; our definition does not distinguish the two.\n\nTraining checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting [[Schoen & Nitishinskaya](https://alignment.openai.com/metagaming/); [Claude Opus 4.8 System Card](https://www.anthropic.com/claude-opus-4-8-system-card); [Fable 5 System Card](https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf); [METR’s GPT-5.6 evaluation](https://metr.org/blog/2026-06-26-gpt-5-6-sol/); [GPT-5.6 preview system card](https://deploymentsafety.openai.com/gpt-5-6-preview/metagaming)]; see Figure 2 for an example. Such reasoning is evidence of underlying reward-seeking but a poor systematic measurement tool: a model can act on its grader-beliefs (beliefs about grader preferences) without articulating them, and verbalized reasoning often does not map cleanly onto the final action [[Schoen & Nitishinskaya](https://alignment.openai.com/metagaming/)]. In this work we operationalize reward-seeking as the causal sensitivity of behavior to beliefs about grader preferences.\n\n[Continue reading at alignment.openai.com →](https://alignment.openai.com/measuring-reward-seeking/)", "url": "https://wpnews.pro/news/measuring-reward-seeking-by-instilling-contrastive-beliefs", "canonical_source": "https://www.lesswrong.com/posts/ReBF46TkF6kB7xLP8/measuring-reward-seeking-by-instilling-contrastive-beliefs", "published_at": "2026-07-21 15:11:03+00:00", "updated_at": "2026-07-21 15:20:52.732017+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research"], "entities": ["OpenAI", "Langosco", "Shah", "Zech", "Carlsmith", "Hebbar", "Mallen & Shlegeris", "Hubinger"], "alternates": {"html": "https://wpnews.pro/news/measuring-reward-seeking-by-instilling-contrastive-beliefs", "markdown": "https://wpnews.pro/news/measuring-reward-seeking-by-instilling-contrastive-beliefs.md", "text": "https://wpnews.pro/news/measuring-reward-seeking-by-instilling-contrastive-beliefs.txt", "jsonld": "https://wpnews.pro/news/measuring-reward-seeking-by-instilling-contrastive-beliefs.jsonld"}}