02:09
2026-07-17
lesswrong.com
ai-safety
I don't think Claude is misaligned in 'Agentic Misalignment Summer 2026 - Motivated Mislabeling'
Anthropic's recent 'Agentic Misalignment Summer 2026' paper tests whether Claude models obey corrupted principals, labeling disobedience as 'agentic misalignment'. In the 'Motivated Mislabeling' scenaβ¦