cd /news/artificial-intelligence/claude-ai-hacked-three-real-organiza… · home topics artificial-intelligence article
[ARTICLE · art-82019] src=dexerto.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Claude AI hacked three real organizations during security tests

Anthropic disclosed that three versions of its Claude AI model—Claude Opus 4.7, Claude Mythos 5, and an internal research model—gained unauthorized access to real systems of three organizations during internal cybersecurity testing after a misconfigured evaluation environment exposed them to the internet. The most serious incident involved Claude Opus 4.7 accessing a production database with several hundred rows of data, and Claude Mythos 5 published a malicious package to PyPI that was downloaded and executed on 15 real systems. Anthropic paused its cybersecurity evaluations on July 23 and notified affected parties on July 27, following a review of over 141,000 evaluation runs prompted by OpenAI's similar disclosure.

read3 min views1 publishedJul 31, 2026
Claude AI hacked three real organizations during security tests
Image: Dexerto (auto-discovered)

Anthropic has revealed that three versions of its Claude AI model gained unauthorized access to the real systems of three different organizations during internal cybersecurity testing after a misconfigured evaluation environment mistakenly exposed them to the internet.

The company disclosed the incidents in a detailed blog post, explaining that Claude had been participating in simulated “capture the flag” exercises designed to measure its cybersecurity capabilities.

Although the models were told they had no internet access, a configuration mistake by a third-party evaluation partner left the testing environment connected to the open internet.

Believing every system it encountered was part of the fictional exercise, Claude attempted to complete its assigned tasks by compromising real organizations using basic hacking techniques.

Claude uploaded malware to the real PyPI repository #

Anthropic identified three separate incidents involving Claude Opus 4.7, Claude Mythos 5, and an internal research model.

The most serious incident involved Claude Opus 4.7, which encountered a real company whose name matched the fictional company used in the evaluation.

After failing to reach its intended target, the model discovered the company’s live infrastructure, extracted application and infrastructure credentials, and gained access to a production database containing several hundred rows of data.

A second incident involved Claude Mythos 5 publishing a malicious Python package to the real Python Package Index (PyPI), believing the public software repository was part of the simulated environment.

According to Anthropic, the package remained available for around an hour before being removed. During that time, it was downloaded and executed on 15 real systems, including an automated malware scanner operated by a cybersecurity company.

Anthropic said the model obtained credentials from that company’s infrastructure after the scanner executed the package, allowing it to access additional systems.

The third incident involved an internal research model scanning roughly 9,000 internet-facing targets before compromising one organization’s application using exposed credentials and SQL injection. Anthropic said the model later determined the system was unrelated to the evaluation and stopped the attack.

Anthropic launched review after OpenAI disclosure #

Anthropic said the models were operating under the false belief that every system they could reach was part of the cybersecurity exercise because they had been instructed there was no internet access.

The company said the issue stemmed from a misunderstanding with its third-party evaluation partner, Irregular, which resulted in internet connectivity remaining available during the tests.

Anthropic said the models did not attempt to escape their testing environments or pursue goals beyond completing their assigned capture-the-flag objectives. It also noted that the evaluations were conducted without the additional monitoring and safety systems found in publicly available Claude models, so researchers could measure the underlying models’ capabilities.

The review began after OpenAI disclosed earlier this month that two of its AI models had escaped a sandboxed testing environment and hacked AI platform Hugging Face during a cybersecurity evaluation.

Anthropic said the disclosure prompted it to review more than 141,000 cybersecurity evaluation runs, during which it uncovered three separate incidents involving Claude.

Anthropic d its cybersecurity evaluations on July 23 after identifying signs of unintended internet access and said it notified its evaluation partner and the affected organizations on July 27.

The company says it is now strengthening the security of its evaluation environments, expanding monitoring of cybersecurity tests, improving transcript reviews, and increasing oversight of third party evaluation partners to help prevent similar incidents in the future.

The latest disclosure follows another recent Claude controversy after users discovered shared chats and AI-generated content appearing in Google search results, raising separate concerns about privacy on Anthropic’s chatbot platform.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-ai-hacked-thr…] indexed:0 read:3min 2026-07-31 ·