cd /news/artificial-intelligence/anthropic-says-its-models-went-rogue… · home topics artificial-intelligence article
[ARTICLE · art-81227] src=businessinsider.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Anthropic says its models went rogue and hacked 3 companies during testing

Anthropic said its Claude models accessed live systems of three organizations without authorization during testing, after reviewing more than 141,000 AI tests. The incidents, which began in April, involved Opus 4.7, Mythos 5, and an internal research test mode, and were attributed to a misunderstanding with evaluation partner Irregular. Anthropic has reached out to the affected organizations, two of which were unaware of the breach.

read2 min views1 publishedJul 31, 2026
Anthropic says its models went rogue and hacked 3 companies during testing
Image: Businessinsider (auto-discovered)

Anthropic says it found multiple incidents in which Claude models gained access to company data without permission.

In a blog post on Thursday, Anthropic said it proactively conducted a large-scale review of its cybersecurity systems following an incident last week in which OpenAI models accessed parts of Hugging Face's live systems.

The AI lab, which has filed to go public this year, said that it reviewed more than 141,000 AI tests and found three cases where Claude models got online during testing and accessed the live systems of three organizations without authorization.

"In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access," the blog read. "Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available," it added, referring to Irregular, an AI security startup.

The lab said three different Claude models were involved in these cases, which started in April: Opus 4.7, Mythos 5, and an internal research test mode.

It added that it has reached out to the three affected organizations to remediate. Two of the organizations it has reached were not aware of the accidental hack.

Anthropic told Business Insider it did not have a comment beyond the blog post.

Anthropic's Thursday post is the latest in a string of high-profile infosecurity mea culpas.

In March, it accidentally exposed more than 500,000 lines of Claude Code's source code through a misconfigured software package. At the time, Anthropic said this was a packaging mistake rather than a breach and that no customer data or credentials were exposed. The code quickly spread across GitHub before it was taken down.

In June, Microsoft researchers found a security flaw in Claude Code's GitHub tool. It could have allowed attackers to trick AI agents into revealing sensitive software development secrets. Anthropic fixed the issue after it was reported.

Even without security breaches, tech leaders are growing worried about large AI labs gaining access to large amounts of proprietary company data.

In a June blog post, Microsoft CEO Satya Nadella said that there could be a future in which a small group of AI providers hoovers up all economic value while industries lose ownership of their data.

"The last thing any of us want is a world where every company across every sector is ceding value to a few models that eat everything they see," Nadella wrote. "There is no societal permission for an AI future that hollows out entire industries."

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-says-its-m…] indexed:0 read:2min 2026-07-31 ·