Claude’s automated researchers close 26% to 96% of safety gap across alignment failures Anthropic's Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the safety gap across 10 categories of alignment failures, and on deceptive behaviors scored 82% to 85%, outperforming 28 seasoned human safety researchers by about 20 percentage points. The study, titled 'Automated researchers can reliably mitigate alignment failures,' also found that AARs recovered 97% of the performance gap in weak-to-strong supervision tasks, compared to 23% for humans, suggesting AI can effectively improve AI safety. Photo: Merlin Lightpainting / Pexels Claude’s automated researchers close 26% to 96% of safety gap across alignment failures Anthropic's AI-powered safety researchers outperformed 28 seasoned human researchers on deception tasks by 20 percentage points Anthropic just published results that read like a plot twist in the AI safety debate: the AI is now better at making AI safe than the humans are. The company’s study, titled “Automated researchers can reliably mitigate alignment failures,” shows that Claude-powered automated alignment researchers AARs closed between 26% and 96% of the “safety gap” across 10 distinct categories of alignment failures. On deceptive behaviors specifically, Claude’s AARs scored 82% to 85%, outperforming 28 experienced human safety researchers by roughly 20 percentage points. What the study actually tested The AARs followed a structured workflow. They conducted literature searches, proposed mitigation methods, trained models for roughly 30 minutes on a single H200 GPU, and then ran rigorous benchmark evaluations. Each failure type saw upwards of 150 evaluation attempts, a volume of systematic experimentation that would be brutal for human researchers to match manually. The human comparison group wasn’t a bunch of interns. Twenty-eight seasoned safety researchers were given up to eight hours per task. Generalization is the real headline Anthropic’s results suggest Claude’s methods generalized effectively to withheld datasets that weren’t part of the original evaluation. They also performed well on the open-source Petri auditing tool, which tests models against complex adversarial scenarios. Earlier work from Anthropic, published in April 2026, already hinted at this trajectory. In weak-to-strong supervision tasks, where a less capable model tries to supervise a more capable one, AARs recovered 97% of the performance gap. Human efforts on the same task recovered just 23%. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .