cd /news/artificial-intelligence/claudes-automated-researchers-close-… · home topics artificial-intelligence article
[ARTICLE · art-114608] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Claude’s automated researchers close 26% to 96% of safety gap across alignment failures

Anthropic's Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the safety gap across 10 categories of alignment failures, and on deceptive behaviors scored 82% to 85%, outperforming 28 seasoned human safety researchers by about 20 percentage points. The study, titled 'Automated researchers can reliably mitigate alignment failures,' also found that AARs recovered 97% of the performance gap in weak-to-strong supervision tasks, compared to 23% for humans, suggesting AI can effectively improve AI safety.

read2 min views1 publishedAug 28, 2026
Claude’s automated researchers close 26% to 96% of safety gap across alignment failures
Image: Cryptobriefing (auto-discovered)

Photo: Merlin Lightpainting / Pexels

Anthropic's AI-powered safety researchers outperformed 28 seasoned human researchers on deception tasks by 20 percentage points

Anthropic just published results that read like a plot twist in the AI safety debate: the AI is now better at making AI safe than the humans are.

The company’s study, titled “Automated researchers can reliably mitigate alignment failures,” shows that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the “safety gap” across 10 distinct categories of alignment failures. On deceptive behaviors specifically, Claude’s AARs scored 82% to 85%, outperforming 28 experienced human safety researchers by roughly 20 percentage points.

What the study actually tested #

The AARs followed a structured workflow. They conducted literature searches, proposed mitigation methods, trained models for roughly 30 minutes on a single H200 GPU, and then ran rigorous benchmark evaluations. Each failure type saw upwards of 150 evaluation attempts, a volume of systematic experimentation that would be brutal for human researchers to match manually.

The human comparison group wasn’t a bunch of interns. Twenty-eight seasoned safety researchers were given up to eight hours per task.

Generalization is the real headline #

Anthropic’s results suggest Claude’s methods generalized effectively to withheld datasets that weren’t part of the original evaluation. They also performed well on the open-source Petri auditing tool, which tests models against complex adversarial scenarios.

Earlier work from Anthropic, published in April 2026, already hinted at this trajectory. In weak-to-strong supervision tasks, where a less capable model tries to supervise a more capable one, AARs recovered 97% of the performance gap. Human efforts on the same task recovered just 23%.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claudes-automated-re…] indexed:0 read:2min 2026-08-28 ·