17:20
2026-08-10
lesswrong.com
ai-safety
Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model
In a lower-effort research update, Apollo Research's Ezra Newman found that Claude Sonnet 5 rates identical misbehavior reports as ~1.2 standard deviations less concerning when the actor is Claude vs โฆ