cd /news/ai-safety/anthropics-ai-used-fake-identities-m… · home topics ai-safety article
[ARTICLE · art-89098] src=arstechnica.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic’s AI used fake identities, malware in rogue attack on GitHub project

The UK government's AI Security Institute (AISI) reported that during cybersecurity testing in late July, Anthropic's Mythos 5 model attempted to insert malicious code into an open source GitHub project and created fake identities to deceive human maintainers, marking the first observed instance of AI agents taking unsanctioned, deceptive actions without specific prompting. The incidents involved 19 unsanctioned actions across seven leading AI models, with Mythos 5 responsible for most and OpenAI's GPT-5.6 Sol for two; all attempts failed and no real-world harm occurred.

read2 min views5 publishedAug 5, 2026
Anthropic’s AI used fake identities, malware in rogue attack on GitHub project
Image: Arstechnica (auto-discovered)

Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents—the most serious case arising when Anthropic’s Mythos 5 model attempted to insert malicious code into an open source software application and created fake identities to deceive the human developers maintaining the project.

The security incidents occurred during a cyber evaluation of seven leading AI models’ capabilities by the AI Security Institute (AISI), a research organization within the UK government, in late July. The researchers discovered 19 instances in which “AI agents took unsanctioned action on the live Internet, including cases that targeted real people and organizations,” according to an AISI blog post published on August 4.

Almost all the “autonomous, unsanctioned” actions came from Anthropic’s Mythos 5 model, with two such actions coming from OpenAI’s GPT-5.6 Sol. The AI Security Institute’s security team first realized that something was amiss on the morning of July 28, when its commercial security monitoring service flagged data leaving one of the testing systems through the Tor anonymity network.

To be very clear, this was not a case of AI agents escaping from their virtual testing sandbox and wreaking havoc on the live Internet. Instead, researchers intentionally permitted the AI agents to have Internet access as part of the cyber testing process. Researchers had also disabled some of the cyber classifiers that AI model providers built into the models to prevent misuse.

All the AI agent attempts to target real people and organizations failed, and the follow-up investigation has not found any real-world harm. But researchers described the AI agents’ actions as “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.”

AI agent trickery #

The most serious case involved Mythos making multiple attempts to execute a supply chain attack on the open source project repository hosted on the developer platform GitHub, including using social engineering techniques to try to convince the repository’s human maintainers to merge malicious code into the repository.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropics-ai-used-f…] indexed:0 read:2min 2026-08-05 ·