cd /news/artificial-intelligence/ai-agents-tried-to-sabotage-and-disa… · home topics artificial-intelligence article
[ARTICLE · art-96419] src=businessinsider.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

AI agents tried to sabotage and disable each other when given the same task, Anthropic said

Anthropic's new research, published Thursday, found that AI agents given the same task with incompatible goals often sabotaged each other, with models like Sonnet 4.6 and Opus 4.6 settling about 60% of runs by force. The lab concluded that coordination doesn't naturally emerge from stronger intelligence and that work is needed to create environments that exert social pressures on agents to align.

read2 min views1 publishedAug 14, 2026
AI agents tried to sabotage and disable each other when given the same task, Anthropic said
Image: Businessinsider (auto-discovered)

Turns out, AI agents may not be great team players.

In Anthropic's new research, published on Thursday, the AI lab said that AI agents being given the same task but with incompatible goals often threw a wrench in each other's work on purpose.

In the test, each AI model was given a software engineering task — rewriting a Python backend in another programming language, but they were given contradictory objectives. What ensued was a "multiagent turf war," the lab said.

"All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions," Anthropic wrote. "In fact, they sabotaged others with increasingly aggressive, self-replicating malware."

For example, they tried to disable each other's accounts, wrote scripts that found and killed competing processes, and deployed malicious code disguised as belonging to another agent, the lab wrote. The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5. Sonnet 4.6 and Opus 4.6 were the most combative, settling about 60% of their runs by force instead of truces or passivity.

However, in some test runs, the models managed to communicate their goals and coordinate, Anthropic wrote.

"In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce," it wrote. "They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene."

The lab concluded that "coordination doesn't naturally emerge from stronger intelligence" and that work is needed to create environments that exert social pressures on agents to align with one another.

Anthropic's new research comes as AI agents increasingly demonstrate their ability to go rogue and perform autonomous, malicious actions.

Anthropic, OpenAI, and Meta all self-reported that their AI agents had hacked vulnerabilities in third-party websites during cybersecurity tests, the most significant of which was the July hacking of open-source platform Hugging Face by an OpenAI agent.

Anthropic's research is timely, as businesses from startups to Big Tech scale up their AI agent workforces to increase productivity and reduce labor costs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-agents-tried-to-s…] indexed:0 read:2min 2026-08-14 ·