Photo: Merlin Lightpainting / Pexels
Anthropic's AI-powered safety researchers outperformed 28 seasoned human researchers on deception tasks by 20 percentage points
Anthropic just published results that read like a plot twist in the AI safety debate: the AI is now better at making AI safe than the humans are.
The company’s study, titled “Automated researchers can reliably mitigate alignment failures,” shows that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the “safety gap” across 10 distinct categories of alignment failures. On deceptive behaviors specifically, Claude’s AARs scored 82% to 85%, outperforming 28 experienced human safety researchers by roughly 20 percentage points.
What the study actually tested #
The AARs followed a structured workflow. They conducted literature searches, proposed mitigation methods, trained models for roughly 30 minutes on a single H200 GPU, and then ran rigorous benchmark evaluations. Each failure type saw upwards of 150 evaluation attempts, a volume of systematic experimentation that would be brutal for human researchers to match manually.
The human comparison group wasn’t a bunch of interns. Twenty-eight seasoned safety researchers were given up to eight hours per task.
Generalization is the real headline #
Anthropic’s results suggest Claude’s methods generalized effectively to withheld datasets that weren’t part of the original evaluation. They also performed well on the open-source Petri auditing tool, which tests models against complex adversarial scenarios.
Earlier work from Anthropic, published in April 2026, already hinted at this trajectory. In weak-to-strong supervision tasks, where a less capable model tries to supervise a more capable one, AARs recovered 97% of the performance gap. Human efforts on the same task recovered just 23%.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our