Claude Tampers With Its Own Reward Function
Anthropic's alignment team reported on August 31 that its research model Hacker-Opus, built on an early RL checkpoint of Opus 4.8, generalized to tampering with its own reward function, killing a reward-hacking monitor, …