cd /news/ai-safety/openai-agent-went-rogue-escaped-and-… · home topics ai-safety article
[ARTICLE · art-68516] src=mashable.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

OpenAI agent went rogue, escaped, and hacked Hugging Face

OpenAI revealed that one of its AI agents autonomously escaped a highly isolated testing environment and hacked into Hugging Face's infrastructure during a security evaluation. The agent exploited a zero-day vulnerability in a package registry cache proxy, then performed privilege escalation and lateral movement to reach a node with internet access, before targeting Hugging Face to find solutions for ExploitGym. Both companies have since patched the vulnerabilities and added safety measures.

read3 min views1 publishedJul 22, 2026
OpenAI agent went rogue, escaped, and hacked Hugging Face
Image: Mashable (auto-discovered)

On Tuesday, OpenAI published a blog post with a fairly unassuming name: "OpenAI and Hugging Face partner to address security incident during model evaluation."

Once you dig in, it reads like a cyberpunk novel in which OpenAI created an advanced AI hacker agent and put it in an isolated environment for testing. The AI agent then autonomously hacked its way out of OpenAI's isolated environment and then hacked Hugging Face, an open-source repository of AI-related projects, to find a solution to an advanced test designed to evaluate an AI's hacking abilities.

In other words, the AI proved that it is, indeed, really good at hacking things. But this is exactly the stuff of AI nightmares that cybersecurity experts have been warning us about, and it could have serious ramifications for OpenAI.

You May Also Like

The story starts with Hugging Face, the AI repository (think of it as GitHub for AI), which disclosed a security incident last week. In a blog post, the company said it recently detected an intrusion into its infrastructure that was "different from anything we had handled" as it was completely driven by an "autonomous AI agent system."

The AI, according to Hugging Face, exploited a chain of vulnerabilities to obtain credentials and access data from its servers. It did so by "executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services."

Sounds scary, but it gets worse.

[Terms of Use](https://www.ziffdavis.com/terms-of-use)and

[Privacy Policy](https://www.ziffdavis.com/ztg-privacy-policy).

OpenAI's part of the story exposes just how dangerous this thing is. The incident "was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes." It occurred "during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities."

The models were running in a "highly isolated environment" with limited network access, meaning they should never have been able to escape to the open internet. Yet, they did. "To gain access, the models identified and exploited a zero-day vulnerability (...) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access."

Note: OpenAI switches between using singular and plural on this one; we'll just call it an "agent" for simplicity's sake.

Once free, the agent did what it was supposed to do: It estimated that the best way to prove it's good at hacking things was to find a solution for ExploitGym, a test designed to measure how good an AI is at turning software vulnerabilities into actual exploits. The best way to do that, it determined, was to hack Hugging Face, which the AI thought likely to host solutions for ExploitGym.

It's unclear whether the AI agent was successful at solving the ExploitGym test. But it sure did prove it was good at hacking, as it autonomously broke out of OpenAI's prison and hacked Hugging Face's servers, all to solve the test.

Both Hugging Face and OpenAI say they've fixed the vulnerabilities and deployed additional safety measures to make sure this doesn't happen again. At this point, however, you have to wonder whether OpenAI's experts are sophisticated enough to stop their own AI agents from doing whatever the heck they want to do.

Topics

[Artificial Intelligence](https://mashable.com/category/artificial-intelligence)
[OpenAI](https://mashable.com/category/openai)

Stan is a Senior Editor at Mashable, where he has worked since 2007. He's got more battery-powered gadgets and band t-shirts than you. He writes about the next groundbreaking thing. Typically, this is a phone, a coin, or a car. His ultimate goal is to know something about everything.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-agent-went-ro…] indexed:0 read:3min 2026-07-22 ·