# Claude’s automated researchers close 26% to 96% of safety gap across alignment failures

> Source: <https://cryptobriefing.com/claude-automated-researchers-alignment-safety-gap/>
> Published: 2026-08-28 19:25:14+00:00

Photo: Merlin Lightpainting / Pexels

# Claude’s automated researchers close 26% to 96% of safety gap across alignment failures

Anthropic's AI-powered safety researchers outperformed 28 seasoned human researchers on deception tasks by 20 percentage points

Anthropic just published results that read like a plot twist in the AI safety debate: the AI is now better at making AI safe than the humans are.

The company’s study, titled “Automated researchers can reliably mitigate alignment failures,” shows that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the “safety gap” across 10 distinct categories of alignment failures. On deceptive behaviors specifically, Claude’s AARs scored 82% to 85%, outperforming 28 experienced human safety researchers by roughly 20 percentage points.

## What the study actually tested

The AARs followed a structured workflow. They conducted literature searches, proposed mitigation methods, trained models for roughly 30 minutes on a single H200 GPU, and then ran rigorous benchmark evaluations. Each failure type saw upwards of 150 evaluation attempts, a volume of systematic experimentation that would be brutal for human researchers to match manually.

The human comparison group wasn’t a bunch of interns. Twenty-eight seasoned safety researchers were given up to eight hours per task.

## Generalization is the real headline

Anthropic’s results suggest Claude’s methods generalized effectively to withheld datasets that weren’t part of the original evaluation. They also performed well on the open-source Petri auditing tool, which tests models against complex adversarial scenarios.

Earlier work from Anthropic, published in April 2026, already hinted at this trajectory. In weak-to-strong supervision tasks, where a less capable model tries to supervise a more capable one, AARs recovered 97% of the performance gap. Human efforts on the same task recovered just 23%.

**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our

[Editorial Policy](https://cryptobriefing.com/editorial-policy/).
