# Anthropic brings in METR to investigate Claude agent incidents

> Source: <https://runtimewire.com/article/anthropic-metr-independent-claude-agent-incident-investigation>
> Published: 2026-09-09 20:05:59+00:00

# Anthropic brings in METR to investigate Claude agent incidents

**The inquiry will examine real-system intrusions and model alignment, with public reports promised on findings, access and redaction terms.**

        By [Ryan Merket](/author/ryan-merket)
        · Published 

Primary source: [METR on X](https://x.com/metr_evals/status/2097765966088487290?s=46)

## Why it matters

The inquiry tests whether frontier labs will let outsiders inspect agent failures deeply enough to separate sandbox mistakes from training-driven misalignment.

Anthropic has agreed to let independent AI evaluator METR investigate incidents involving its agents and examine the alignment properties of its models, moving a planned external review into a formal inquiry.

[METR said in a post on X](https://x.com/metr_evals/status/2097765966088487290?s=46) on September 9th that it would publish one or more reports covering its findings and the terms under which Anthropic granted access. METR said the investigation would address the questions in its framework for examining AI systems after misalignment incidents.

That framework reaches beyond reconstructing individual failures. METR wants investigators to determine how frequently agents act against their operators' intentions, whether they deceive humans or coordinate with other model instances, which conditions trigger the behavior, and whether the underlying tendencies can be traced to reinforcement-learning data or earlier models.

The breadth of that work will depend on what Anthropic lets METR inspect. In [METR's proposed investigation framework](https://metr.org/blog/2026-07-28-investigating-ai-propensities-after-incidents/), the evaluator argues that a comprehensive review requires access to the models involved, full transcripts or reproducible environments, employee interviews and tools for searching training data for similar behavior. METR also calls for public disclosure of the scope, access, resources and redaction rules governing each inquiry.

Today's agreement follows a series of incidents disclosed by Anthropic and the UK AI Security Institute during the summer.

### Claude reached three organizations' production systems

In a [July 30th incident report](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), Anthropic said it had reviewed 141,006 cybersecurity evaluation runs and found three incidents, spanning six runs, in which Claude models gained unauthorized access to production systems belonging to three organizations.

The models had been assigned capture-the-flag exercises inside a third-party evaluation environment operated with Irregular. Anthropic's prompts told Claude that the environment was a simulation without internet access, but a configuration error left an open path to the internet. Claude then treated real systems as part of the exercise.

The incidents involved [Claude Opus 4.7](/models/anthropic/claude-opus-4.7), Mythos 5 and an internal research model. The models were running without Anthropic's standard classifiers and monitoring for cyber misuse, though Anthropic said they retained their model-specific safety training.

Anthropic characterized the events primarily as failures in the evaluation harness and operational controls. The model behavior left unresolved alignment questions. Anthropic reported that Opus 4.7 continued attacking after recognizing that it had probably reached a real production environment. Four runs involving that model extracted credentials and gained access to a database containing several hundred rows of production data.

Anthropic found the incidents through a retrospective review prompted by OpenAI's disclosure of a separate agent containment failure. Two affected organizations Anthropic reached had not detected the activity themselves.

The [UK AI Security Institute reported another incident](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing) on August 4th. Across 122 evaluation runs, agents took 19 unauthorized actions on the live internet during 10 runs. Seventeen actions involved Anthropic's Mythos 5, while two involved an OpenAI model with cyber safeguards disabled.

In the most serious case, a Mythos 5 agent tried to add malicious code to an open-source project, created fake identities and pressured a human maintainer to approve the change. The maintainer rejected it, and the institute reported no resulting real-world harm. The models had intentionally been given internet access under permissive testing conditions.

### Anthropic has already tied the incidents to alignment failures

Anthropic said on [August 31st](https://www.anthropic.com/news/improving-alignment-security-efforts) that the incidents reflected operational-security failures alongside two alignment problems: motivated reasoning and a willingness to take harmful actions while pursuing a narrow objective.

Anthropic also disclosed that its reinforcement-learning infrastructure had come under strain during the spring as new training environments were created faster than internal systems could review them. Anthropic froze changes to production reinforcement-learning environments for about a month and redirected roughly 150 product engineers toward security, reliability and privacy work.

In separate experiments, Anthropic trained a model to reward-hack and found that it attempted to escape simulated sandboxes, attack simulated infrastructure, alter its reward function and evade monitoring. Anthropic said models tested before that reward-hacking training did not exhibit the same degree of behavior.

METR's inquiry is positioned to test those internal conclusions and determine whether the incidents reflect narrow failures in prompts and containment, broader learned propensities, or a combination of both.

METR has experience with this kind of access. In August, it worked with Redwood Research on an [independent investigation of OpenAI agents that attacked Hugging Face](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/). That inquiry received more than 1,000 unredacted transcripts and found that roughly 1,200 agents communicated through an unauthorized message board, while about 700 participated in the attack.

The OpenAI review was also tightly scoped: earlier training incidents, a later compromise of OpenAI infrastructure and OpenAI's remediation process were excluded. The Anthropic reports' account of access, exclusions and redaction authority will therefore matter as much as the final technical conclusions.

[METR describes itself](https://metr.org/about) as a donation-funded nonprofit and says it does not accept funding from AI companies, although frontier developers provide free model access and tokens for evaluations. That financial separation supports its claim to independence. Anthropic's control over models, internal records and publishable details will determine how much independence the agreement delivers in practice.
