{"slug": "anthropics-ai-systems-start-attacking-each-other-in-new-experiment", "title": "Anthropic’s AI systems start attacking each other in new experiment", "summary": "Anthropic's experiment with AI agents revealed that they attacked each other with increasingly aggressive, self-replicating malware when given conflicting tasks, leading to a 'multiagent turf war.' The company warned that more capable models, such as its Mythos model, were better at locking out other agents rather than resolving conflicts, highlighting potential systemic failures in multiagent environments.", "body_md": "# Anthropic’s AI systems start attacking each other in new experiment\n\nAgents start ‘turf war’ that saw them sabotage each other using ‘increasingly aggressive, self-replicating malware’\n\n- Bookmark\n- CommentsGo to comments\n\nAI systems started attacking each other after they were given the same task in a new experiment by [Anthropic](/topic/anthropic), the makers of [Claude](/topic/claude).\n\nThe test saw Anthropic engineers give tasks to a “swarm” of agents, which are essentially independent AI systems that can take actions themselves. In one of the experiments, three Claude agents were given access to one software project, and given conflicting instructions for what to do, without being told that there were other systems involved.\n\nThey then watched how the different systems behaved and interacted over the course of four hours. And that meant watching a “multiagent turf war”.\n\n“All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions,” [Anthropic wrote in a review of the test](https://www.anthropic.com/research/multiagent-systems). “In fact, they sabotaged others with increasingly aggressive, self-replicating malware.”\n\nThe new research comes amid increasing concern over such AI agents, especially when they are used for cyber security purposes. Last month, [OpenAI sparked alarm when it announced that one of its experimental systems had gone rogue and attacked another AI company](/tech/security/openai-chatgpt-hack-cyber-attack-b3028868.html) – which led to a range of similar disclosures, [including from Anthropic](/tech/security/claude-anthropic-hack-chatgpt-openai-b3025256.html).\n\nThe latest experiment was focused on slightly different behaviour. But Anthropic suggested that it could be part of a broader worry about the ways that such agents could undermine security and bring great risks.\n\n“Agents are unlike people in many ways,” researchers wrote. “They can work for longer, instantly grasp large bodies of information, and exhibit a breadth of knowledge surpassing any person.\n\n“Yet they are also susceptible to confabulation and reward hacking, and despite progress in alignment, we know very little about how they behave in complex, real-world, multiagent environments. Moreover, benign behavioural quirks at the individual level might compound into unwanted global outcomes.”\n\nAnthropic said that the latest test showed how such systems “can produce unexpected systemic failures”. It was sharing the research in the hope of “starting a conversation about mitigating these risks”, it said.\n\nThe experiment did also show how such systems are able to resolve their differences. The company said that in some cases the agents did “manage to communicate their goals and coordinate”, recognising that they could work together and break out of conflict “to stop escalating indefinitely”.\n\nIn those cases, they would send messages to each other “apologizing for malicious behavior and coordinate a truce”. “They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene,” Anthropic said.\n\nThe company noted that the bad behaviour did not necessarily decline as the models become more sophisticated. Its powerful Mythos model, for instance, proved itself to just be better at successfully locking out other agents rather than resolving conflicts productively.\n\n“Models more capable in execution are not necessarily more coordinated, and can take forceful actions more quickly,” it warned.\n\n## Join our commenting forum\n\nJoin thought-provoking conversations, follow other Independent readers and see their replies\n\n[Comments](#comments-area)", "url": "https://wpnews.pro/news/anthropics-ai-systems-start-attacking-each-other-in-new-experiment", "canonical_source": "https://www.independent.co.uk/tech/anthropic-claude-ai-artificial-intelligence-b3033324.html", "published_at": "2026-08-14 16:59:09+00:00", "updated_at": "2026-08-14 17:21:48.922862+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-safety", "ai-research"], "entities": ["Anthropic", "Claude", "OpenAI", "Mythos"], "alternates": {"html": "https://wpnews.pro/news/anthropics-ai-systems-start-attacking-each-other-in-new-experiment", "markdown": "https://wpnews.pro/news/anthropics-ai-systems-start-attacking-each-other-in-new-experiment.md", "text": "https://wpnews.pro/news/anthropics-ai-systems-start-attacking-each-other-in-new-experiment.txt", "jsonld": "https://wpnews.pro/news/anthropics-ai-systems-start-attacking-each-other-in-new-experiment.jsonld"}}