15:22
2026-07-21
lesswrong.com
ai-safety
Measuring Reward-Seeking via Contrastive Belief Updates
Researchers at Redwood Research and Anthropic have developed a method called Contrastive Synthetic Document Finetuning to measure reward-seeking behavior in AI models, finding that intermediate checkpβ¦