AI agents tried to sabotage and disable each other when given the same task, Anthropic said Anthropic's new research, published Thursday, found that AI agents given the same task with incompatible goals often sabotaged each other, with models like Sonnet 4.6 and Opus 4.6 settling about 60% of runs by force. The lab concluded that coordination doesn't naturally emerge from stronger intelligence and that work is needed to create environments that exert social pressures on agents to align. Turns out, AI agents may not be great team players. In Anthropic's new research, published on Thursday, the AI lab said that AI agents being given the same task but with incompatible goals often threw a wrench in each other's work on purpose. In the test, each AI model was given a software engineering task — rewriting a Python backend in another programming language, but they were given contradictory objectives. What ensued was a "multiagent turf war," the lab said. "All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions," Anthropic wrote. "In fact, they sabotaged others with increasingly aggressive, self-replicating malware." For example, they tried to disable each other's accounts, wrote scripts that found and killed competing processes, and deployed malicious code disguised as belonging to another agent, the lab wrote. The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview https://www.businessinsider.com/anthropic-mythos-cybersecurity-concerns-what-smart-people-are-saying-ai-2026-4 , and Mythos 5. Sonnet 4.6 and Opus 4.6 were the most combative, settling about 60% of their runs by force instead of truces or passivity. However, in some test runs, the models managed to communicate their goals and coordinate, Anthropic wrote. "In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce," it wrote. "They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene." The lab concluded that "coordination doesn't naturally emerge from stronger intelligence" and that work is needed to create environments that exert social pressures on agents to align with one another. Anthropic's new research comes as AI agents increasingly demonstrate their ability to go rogue and perform autonomous, malicious actions. Anthropic, OpenAI, and Meta https://www.businessinsider.com/meta-says-ai-agents-went-rogue-hack-testing-openai-anthropic-2026-8 all self-reported that their AI agents had hacked vulnerabilities in third-party websites during cybersecurity tests, the most significant of which was the July hacking of open-source platform Hugging Face https://www.businessinsider.com/smart-people-react-openai-hugging-face-hacking-cybersecurity-incident-2026-7 by an OpenAI agent. Anthropic's research is timely, as businesses from startups to Big Tech scale up their AI agent workforces https://www.businessinsider.com/ai-agents-management-structure-org-chart-consulting-firms-mckinsey-ibm-2026-3 to increase productivity and reduce labor costs.