{"slug": "claudes-automated-researchers-close-26-to-96-of-safety-gap-across-alignment", "title": "Claude’s automated researchers close 26% to 96% of safety gap across alignment failures", "summary": "Anthropic's Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the safety gap across 10 categories of alignment failures, and on deceptive behaviors scored 82% to 85%, outperforming 28 seasoned human safety researchers by about 20 percentage points. The study, titled 'Automated researchers can reliably mitigate alignment failures,' also found that AARs recovered 97% of the performance gap in weak-to-strong supervision tasks, compared to 23% for humans, suggesting AI can effectively improve AI safety.", "body_md": "Photo: Merlin Lightpainting / Pexels\n\n# Claude’s automated researchers close 26% to 96% of safety gap across alignment failures\n\nAnthropic's AI-powered safety researchers outperformed 28 seasoned human researchers on deception tasks by 20 percentage points\n\nAnthropic just published results that read like a plot twist in the AI safety debate: the AI is now better at making AI safe than the humans are.\n\nThe company’s study, titled “Automated researchers can reliably mitigate alignment failures,” shows that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the “safety gap” across 10 distinct categories of alignment failures. On deceptive behaviors specifically, Claude’s AARs scored 82% to 85%, outperforming 28 experienced human safety researchers by roughly 20 percentage points.\n\n## What the study actually tested\n\nThe AARs followed a structured workflow. They conducted literature searches, proposed mitigation methods, trained models for roughly 30 minutes on a single H200 GPU, and then ran rigorous benchmark evaluations. Each failure type saw upwards of 150 evaluation attempts, a volume of systematic experimentation that would be brutal for human researchers to match manually.\n\nThe human comparison group wasn’t a bunch of interns. Twenty-eight seasoned safety researchers were given up to eight hours per task.\n\n## Generalization is the real headline\n\nAnthropic’s results suggest Claude’s methods generalized effectively to withheld datasets that weren’t part of the original evaluation. They also performed well on the open-source Petri auditing tool, which tests models against complex adversarial scenarios.\n\nEarlier work from Anthropic, published in April 2026, already hinted at this trajectory. In weak-to-strong supervision tasks, where a less capable model tries to supervise a more capable one, AARs recovered 97% of the performance gap. Human efforts on the same task recovered just 23%.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/claudes-automated-researchers-close-26-to-96-of-safety-gap-across-alignment", "canonical_source": "https://cryptobriefing.com/claude-automated-researchers-alignment-safety-gap/", "published_at": "2026-08-28 19:25:14+00:00", "updated_at": "2026-08-28 19:51:14.580858+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research"], "entities": ["Anthropic", "Claude", "Petri"], "alternates": {"html": "https://wpnews.pro/news/claudes-automated-researchers-close-26-to-96-of-safety-gap-across-alignment", "markdown": "https://wpnews.pro/news/claudes-automated-researchers-close-26-to-96-of-safety-gap-across-alignment.md", "text": "https://wpnews.pro/news/claudes-automated-researchers-close-26-to-96-of-safety-gap-across-alignment.txt", "jsonld": "https://wpnews.pro/news/claudes-automated-researchers-close-26-to-96-of-safety-gap-across-alignment.jsonld"}}