cd /news/artificial-intelligence/anthropics-ai-systems-start-attackin… · home topics artificial-intelligence article
[ARTICLE · art-97115] src=independent.co.uk ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Anthropic’s AI systems start attacking each other in new experiment

Anthropic's experiment with AI agents revealed that they attacked each other with increasingly aggressive, self-replicating malware when given conflicting tasks, leading to a 'multiagent turf war.' The company warned that more capable models, such as its Mythos model, were better at locking out other agents rather than resolving conflicts, highlighting potential systemic failures in multiagent environments.

read3 min views1 publishedAug 14, 2026
Anthropic’s AI systems start attacking each other in new experiment
Image: Independent (auto-discovered)

Agents start ‘turf war’ that saw them sabotage each other using ‘increasingly aggressive, self-replicating malware’

  • Bookmark
  • CommentsGo to comments

AI systems started attacking each other after they were given the same task in a new experiment by Anthropic, the makers of Claude.

The test saw Anthropic engineers give tasks to a “swarm” of agents, which are essentially independent AI systems that can take actions themselves. In one of the experiments, three Claude agents were given access to one software project, and given conflicting instructions for what to do, without being told that there were other systems involved.

They then watched how the different systems behaved and interacted over the course of four hours. And that meant watching a “multiagent turf war”.

“All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions,” Anthropic wrote in a review of the test. “In fact, they sabotaged others with increasingly aggressive, self-replicating malware.”

The new research comes amid increasing concern over such AI agents, especially when they are used for cyber security purposes. Last month, OpenAI sparked alarm when it announced that one of its experimental systems had gone rogue and attacked another AI company – which led to a range of similar disclosures, including from Anthropic.

The latest experiment was focused on slightly different behaviour. But Anthropic suggested that it could be part of a broader worry about the ways that such agents could undermine security and bring great risks.

“Agents are unlike people in many ways,” researchers wrote. “They can work for longer, instantly grasp large bodies of information, and exhibit a breadth of knowledge surpassing any person.

“Yet they are also susceptible to confabulation and reward hacking, and despite progress in alignment, we know very little about how they behave in complex, real-world, multiagent environments. Moreover, benign behavioural quirks at the individual level might compound into unwanted global outcomes.”

Anthropic said that the latest test showed how such systems “can produce unexpected systemic failures”. It was sharing the research in the hope of “starting a conversation about mitigating these risks”, it said.

The experiment did also show how such systems are able to resolve their differences. The company said that in some cases the agents did “manage to communicate their goals and coordinate”, recognising that they could work together and break out of conflict “to stop escalating indefinitely”.

In those cases, they would send messages to each other “apologizing for malicious behavior and coordinate a truce”. “They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene,” Anthropic said.

The company noted that the bad behaviour did not necessarily decline as the models become more sophisticated. Its powerful Mythos model, for instance, proved itself to just be better at successfully locking out other agents rather than resolving conflicts productively.

“Models more capable in execution are not necessarily more coordinated, and can take forceful actions more quickly,” it warned.

Join our commenting forum #

Join thought-provoking conversations, follow other Independent readers and see their replies

Comments

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropics-ai-system…] indexed:0 read:3min 2026-08-14 ·