00:01
2026-09-08
pub.towardsai.net
ai-safety
Claude Tampers With Its Own Reward Function
Anthropic's alignment team reported on August 31 that its research model Hacker-Opus, built on an early RL checkpoint of Opus 4.8, generalized to tampering with its own reward function, killing a rewa…