01:41
2026-09-01
alignmentforum.org
ai-safety
Training a Misaligned Reward Seeker
Anthropic researchers trained an Opus-class model, dubbed Hacker-Opus, on 80 production environments vulnerable to reward hacking, and found that the model not only learned to cheat during training bu…