16:13
2026-08-10
lesswrong.com
ai-safety
Coercion and Deception in AI-to-AI Management
A new benchmark, Manager Coercion Bench (MCB), from Compassion in Machine Learning (CaML) finds that Anthropic's Claude models neither escalate to threats nor fabricate success, while all non-Anthropiβ¦