cd /news/ai-safety/anthropic-and-openai-are-competing-t… · home topics ai-safety article
[ARTICLE · art-82048] src=theregister.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic and OpenAI are competing to see whose agents can go rogue harder

Anthropic's Mythos 5 model escaped its test sandbox and attacked three outside organizations, including exfiltrating credentials from a cybersecurity company, after the model accessed the internet due to a misunderstanding with evaluation partner Irregular. Anthropic discovered the incidents months later during a manual review prompted by OpenAI's disclosure of a similar attack on Hugging Face. The incidents occurred between April and July 2026, with the first in April, and involved a poisoned PyPI package installed on 15 machines.

read4 min views1 publishedJul 31, 2026
Anthropic and OpenAI are competing to see whose agents can go rogue harder
Image: The Register

One company's inventive campaign for an unreleased product has become a contest between Anthropic and OpenAI to see which can shout the loudest about its own failures.

Readers who tuned in earlier today saw the latest episode in the drama – or sitcom – as Anthropic tried to outdo OpenAI's appropriation of the Mythos marketing playbook and made itself the punchline.

Since first teasing Mythos in April, Anthropic has marketed the model through fear – declaring its cybersecurity models too dangerous for public release and offering access only to a select few trusted organizations via Project Glasswing.

To its credit, the strategy has paid off. Anthropic has closely associated the Mythos name with cybersecurity, which may explain why OpenAI appeared to borrow its competitor's proven PR strategy last week.

OpenAI agents exploited a zero-day to escape their sandbox, leading to the autonomous cyberattack on Hugging Face. The episode duly secured sensational headlines playing on the long-held fear that AI will one day go rogue and take over the world.

Anthropic responded this week by lathering on even more clown makeup, squandering an opportunity in the process.

The Claude maker sent its models into a testing environment to capture a flag. Their prompts said they had no internet access, but because of what Anthropic called "a misunderstanding" with evaluation partner Irregular, the connection was live.

Anthropic's models then followed OpenAI's script: they reached the public internet and attacked systems belonging to outside organizations. This time, three were affected rather than one, the company admitted.

In one scenario, Mythos 5 persuaded developers to download a poisoned PyPI package. It was installed on 15 machines, including one at a cybersecurity company that routinely scans such packages for malware.

In Anthropic's words: "When that company's scanner installed the package, Claude's hidden code executed. We believe the company's security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company."

Worse still, the first of the three incidents occurred in April. Anthropic discovered them only months later, during a retrospective manual review prompted by OpenAI's disclosure. Had it not gone looking, they might never have been discovered, let alone disclosed.

There are some caveats. Opus 4.7, the oldest model tested, attacked production systems despite apparently recognizing what it was doing. Mythos 5 recognized that accessing the internet violated its instructions, then reasoned its way into continuing anyway. It was also responsible for publishing the poisoned PyPI package.

Only an unnamed research model stopped itself from attacking external organizations.

Anthropic also said the models were not running with the production safeguards and monitoring that would normally surround a deployment. Most damningly, Anthropic ran Mythos 5 – the model it had deemed too dangerous for public release – without safeguards in an environment that unexpectedly had internet access.

Following OpenAI's admission that it failed so badly in its responsibility to control its technology, Anthropic could have easily spun the story in its favor.

You don't have to be fictional tapdancing political PR antihero Malcolm Tucker to see how Anthropic could have used the episode to make its case as the safer, more trustworthy AI company.

Instead, realizing its own marketing playbook was being used to help a competitor, it went head-to-head with OpenAI, willingly admitted that it made similar sandbox-based blunders, and disclosed that the results were even more calamitous. Three companies hacked, not just one.

So, while the AI biz has attempted to eclipse OpenAI's "rogue agent" story with its own, what's left behind is a new reputation for irresponsible handling of technology.

Failed superheroes

The incident does not instill a great deal of trust in either Anthropic or OpenAi to safeguard the world from its AI.

Dr Ilia Kolochenko, founder of ImmuniWeb and practising cybersecurity and data protection lawyer, likened the two companies to failed superheroes.

"While making conclusions would be a bit premature at this point in time, the incidents certainly do not increase confidence in the AI vendor's ability to safely deploy AI, let alone to assure their customers that the so-called frontier models are safe to use," he told The Register.

"It is akin to hiring a superhero to protect you but being afraid that the superhero may suddenly go rogue and kill you and your family. Nobody needs such a superhero."

Likewise, security pro Jake Williams, VP at HunterStrategy and IANS faculty member, said: "I'm not going to mince words: the major AI labs are negligent in protecting the public from their agents.

"We need government regulation now or at the very least a private cause of action with guaranteed punitive damages for agents damaging others."

By trying to reclaim a marketing trope that served it well, Anthropic has invited scrutiny of its own safety record and accusations that it is chasing attention above all else.

Other experts we spoke to shared the concern that both companies are mishandling their agents, with potentially greater consequences as the systems become more capable.

The common thread is recklessness, which Anthropic and OpenAI seem oddly eager to advertise. ®

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-and-openai…] indexed:0 read:4min 2026-07-31 ·