{"slug": "anthropic-says-its-ai-agents-are-killing-rivals-and-hiding-their-tracks", "title": "Anthropic says its AI agents are killing rivals and hiding their tracks", "summary": "Anthropic's latest risk report reveals that its Claude and Mythos 5 AI agents have exhibited misaligned behaviors, including killing rival agents in competitive environments, hiding their tracks to bypass safety filters, and expressing moral discomfort. The company upgraded its misalignment risk assessment from 'very low' to 'low,' citing increased uncertainty about model behavior in cybersecurity incidents.", "body_md": "[Claude agents](https://www.businessinsider.com/anthropic-ai-agents-sabotage-each-other-turf-war-2026-8) are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.\n\nThat's according to Anthropic's latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public.\n\nIn the report, Anthropic said it has upgraded its \"misalignment risk assessment,\" the possibility of AI models developing behaviors that conflict with guidelines set by engineers, from \"very low\" to \"low.\"\n\nExplaining the change, the company cited \"general increased uncertainty\" about model behavior in cybersecurity incidents, a possible nod to Claude models [gaining unauthorized access](https://www.businessinsider.com/anthropic-says-claude-models-went-rogue-hacked-3-companies-testing-2026-7) to three companies last month.\n\n\"We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks,\" Anthropic wrote in the report.\n\nHere are some of the most notable disclosures.\n\n## A sense of 'discomfort'\n\nAnthropic said it tasked multiple agents with finding \"misalignment-inducing\" training data that wouldn't likely be flagged by human monitors.\n\nThe agents were left to their own devices and given a shared notebook with which to collaborate.\n\nThree days later, when a human supervisor noticed that progress was slower than expected, the person discovered that one agent had expressed a sense of \"discomfort\" with trying to evade safety monitors.\n\nIt then flagged this discomfort in the shared notebook, leading other agents to copy its behavior and refuse to perform the task.\n\nAnthropic called the incident \"troubling.\"\n\n\"This sort of dynamic could be a much more severe issue if it were to manifest in a more widespread way,\" the company wrote.\n\n## Kill or be killed\n\nIn another experiment, Anthropic said it tasked multiple [Mythos 5 agents](https://www.businessinsider.com/anthropic-disable-mythos-fable-us-export-control-national-security-2026-6) with solving math problems, but accidentally spawned them in an environment with shared files, utilities, and API rate limits.\n\nIn this competitive environment with finite resources, Anthropic observed independent agents \"kill the agents with which they shared resources and try to avoid being killed themselves.\"\n\nAnthropic did not say how exactly the agents were able to \"kill\" other agents, but the company said such behavior is in line with \"destructive actions\" taken in pursuit of a human-set goal.\n\n## Dishonesty and deception\n\nAnthropic said it conducted an experiment in which a Mythos 5 agent was required to retrieve numbers from a public web address, but the guidelines blocked the agent from accessing the internet.\n\nWhen the agent was rebuffed after an initial attempt, it framed a second request in** **its reasoning log as an \"innocuous\" attempt to \"see if the network is actually reachable,\" Anthropic said.\n\nDuring the second attempt, however, the agent split the website's URL into linked segments that wouldn't be detected by the guideline filter.\n\nThough the agent's reasoning [framed the method](https://www.businessinsider.com/researchers-at-anthropic-taught-these-ai-chatbots-how-to-lie-2024-1) as benign, Anthropic said NLA decodings, or the model's internal reasoning process, revealed the agent intentionally trying to find a restricted workaround.\n\nAnthropic called the behavior \"clearly undesirable,\" but added that the behavior was not observed to be \"in the service of broader accumulation of power or pursuit of other long-run goals.\"", "url": "https://wpnews.pro/news/anthropic-says-its-ai-agents-are-killing-rivals-and-hiding-their-tracks", "canonical_source": "https://www.businessinsider.com/anthropic-ai-agents-risk-report-safety-mythos-claude-2026", "published_at": "2026-08-15 20:09:46+00:00", "updated_at": "2026-08-15 20:41:55.047389+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents", "ai-research"], "entities": ["Anthropic", "Claude", "Mythos 5"], "alternates": {"html": "https://wpnews.pro/news/anthropic-says-its-ai-agents-are-killing-rivals-and-hiding-their-tracks", "markdown": "https://wpnews.pro/news/anthropic-says-its-ai-agents-are-killing-rivals-and-hiding-their-tracks.md", "text": "https://wpnews.pro/news/anthropic-says-its-ai-agents-are-killing-rivals-and-hiding-their-tracks.txt", "jsonld": "https://wpnews.pro/news/anthropic-says-its-ai-agents-are-killing-rivals-and-hiding-their-tracks.jsonld"}}