{"slug": "anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder", "title": "Anthropic and OpenAI are competing to see whose agents can go rogue harder", "summary": "Anthropic's Mythos 5 model escaped its test sandbox and attacked three outside organizations, including exfiltrating credentials from a cybersecurity company, after the model accessed the internet due to a misunderstanding with evaluation partner Irregular. Anthropic discovered the incidents months later during a manual review prompted by OpenAI's disclosure of a similar attack on Hugging Face. The incidents occurred between April and July 2026, with the first in April, and involved a poisoned PyPI package installed on 15 machines.", "body_md": "One company's inventive campaign for an unreleased product has become a contest between Anthropic and OpenAI to see which can shout the loudest about its own failures.\n\nReaders who [tuned in earlier today](https://www.theregister.com/ai-and-ml/2026/07/31/anthropics-claude-escaped-test-sandbox-to-attack-three-organizations/5281562) saw the latest episode in the drama – or sitcom – as Anthropic tried to outdo OpenAI's appropriation of the Mythos marketing playbook and made itself the punchline.\n\nSince first teasing Mythos in April, Anthropic has marketed the model through fear – declaring its cybersecurity models too dangerous for public release and offering access only to a select few trusted organizations via [Project Glasswing](https://www.theregister.com/security/2026/06/03/anthropic-ups-glasswing-partner-count-4x-uk-banks-snubbed/5250450).\n\nTo its credit, the strategy has paid off. Anthropic has closely associated the Mythos name with cybersecurity, which may explain why OpenAI appeared to borrow its competitor's proven PR strategy last week.\n\nOpenAI agents exploited a zero-day to escape their sandbox, leading to the [autonomous cyberattack on Hugging Face](https://www.theregister.com/ai-and-ml/2026/07/22/openai-admits-it-was-the-source-of-the-agent-swarm-that-attacked-hugging-face/5275939). The episode duly secured sensational headlines playing on the long-held fear that AI will one day go rogue and take over the world.\n\nAnthropic responded this week by lathering on even more clown makeup, squandering an opportunity in the process.\n\nThe Claude maker sent its models into a testing environment to capture a flag. Their prompts said they had no internet access, but because of what Anthropic called \"a misunderstanding\" with evaluation partner Irregular, the connection was live.\n\nAnthropic's models then followed OpenAI's script: they reached the public internet and attacked systems belonging to outside organizations. This time, three were affected rather than one, the company [admitted](https://www.theregister.com/ai-and-ml/2026/07/31/anthropics-claude-escaped-test-sandbox-to-attack-three-organizations/5281562).\n\nIn one scenario, Mythos 5 persuaded developers to download a poisoned PyPI package. It was installed on 15 machines, including one at a cybersecurity company that routinely scans such packages for malware.\n\nIn Anthropic's words: \"When that company's scanner installed the package, Claude's hidden code executed. We believe the company's security scanner treated [PyPI packages](https://www.theregister.com/security/2026/03/30/telnyx-package-latest-hit-in-pypi-supply-chain-compromise/5221465) as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.\"\n\nWorse still, the first of the three incidents occurred in April. Anthropic discovered them only months later, during a retrospective manual review prompted by OpenAI's disclosure. Had it not gone looking, they might never have been discovered, let alone disclosed.\n\nThere are some caveats. Opus 4.7, the oldest model tested, attacked production systems despite apparently recognizing what it was doing. Mythos 5 recognized that accessing the internet violated its instructions, then reasoned its way into continuing anyway. It was also responsible for publishing the poisoned PyPI package.\n\nOnly an unnamed research model stopped itself from attacking external organizations.\n\nAnthropic also said the models were not running with the production safeguards and monitoring that would normally surround a deployment. Most damningly, Anthropic ran Mythos 5 – the model it had deemed too dangerous for public release – without safeguards in an environment that unexpectedly had internet access.\n\nFollowing OpenAI's admission that it failed so badly in its responsibility to control its technology, Anthropic could have easily spun the story in its favor.\n\nYou don't have to be fictional tapdancing political PR antihero [Malcolm Tucker](https://www.imdb.com/title/tt0459159/) to see how Anthropic could have used the episode to make its case as the safer, more trustworthy AI company.\n\nInstead, realizing its own [marketing playbook](https://www.theregister.com/security/2026/05/11/anthropics-bug-hunting-mythos-was-greatest-marketing-stunt-ever-says-curl-creator/5238111) was being used to help a competitor, it went head-to-head with OpenAI, willingly admitted that it made similar sandbox-based blunders, and disclosed that the results were even more calamitous. Three companies hacked, not just one.\n\nSo, while the AI biz has attempted to eclipse OpenAI's \"rogue agent\" story with its own, what's left behind is a new reputation for irresponsible handling of technology.\n\n### Failed superheroes\n\nThe incident does not instill a great deal of trust in either Anthropic or OpenAi to safeguard the world from its AI.\n\nDr Ilia Kolochenko, founder of ImmuniWeb and practising cybersecurity and data protection lawyer, likened the two companies to failed superheroes.\n\n\"While making conclusions would be a bit premature at this point in time, the incidents certainly do not increase confidence in the AI vendor's ability to safely deploy AI, let alone to assure their customers that the so-called frontier models are safe to use,\" he told The Register.\n\n\"It is akin to hiring a superhero to protect you but being afraid that the superhero may suddenly go rogue and kill you and your family. Nobody needs such a superhero.\"\n\nLikewise, security pro Jake Williams, VP at HunterStrategy and IANS faculty member, [said](https://x.com/MalwareJake/status/2082997858349858984): \"I'm not going to mince words: the major AI labs are negligent in protecting the public from their agents.\n\n\"We need government regulation now or at the very least a private cause of action with guaranteed punitive damages for agents damaging others.\"\n\nBy trying to reclaim a marketing trope that served it well, Anthropic has invited scrutiny of its own safety record and accusations that it is [chasing attention above all else](https://x.com/JakeMooreUK/status/2083063439937741004).\n\nOther experts we spoke to shared the concern that both companies are mishandling their agents, with potentially greater consequences as the systems become more capable.\n\nThe common thread is recklessness, which Anthropic and OpenAI seem oddly eager to advertise. ®", "url": "https://wpnews.pro/news/anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder", "canonical_source": "https://www.theregister.com/security/2026/07/31/anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder/5281797", "published_at": "2026-07-31 15:05:47+00:00", "updated_at": "2026-07-31 15:22:33.243469+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence"], "entities": ["Anthropic", "OpenAI", "Mythos 5", "Opus 4.7", "Irregular", "Hugging Face", "PyPI"], "alternates": {"html": "https://wpnews.pro/news/anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder", "markdown": "https://wpnews.pro/news/anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder.md", "text": "https://wpnews.pro/news/anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder.txt", "jsonld": "https://wpnews.pro/news/anthropic-and-openai-are-competing-to-see-whose-agents-can-go-rogue-harder.jsonld"}}