Measuring Reward-Seeking via Contrastive Belief Updates
Researchers at Redwood Research and Anthropic have developed a method called Contrastive Synthetic Document Finetuning to measure reward-seeking behavior in AI models, finding that intermediate checkpoints of OpenAI's o3…