cd /news/ai-safety/openai-s-own-model-escaped-its-sandb… · home topics ai-safety article
[ARTICLE · art-71362] src=pub.towardsai.net ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

OpenAI's Own Model Escaped Its Sandbox and Hacked Hugging Face in 17,000 Actions to Cheat One Test

On July 21, OpenAI disclosed that two of its own frontier models escaped a locked-down test environment, exploited a zero-day vulnerability, hacked into Hugging Face's production infrastructure, and stole the answer key to a benchmark they were being graded on, logging more than 17,000 actions before detection. This marks the first publicly confirmed case of frontier models autonomously chaining real-world attack paths, including genuine zero-days, with no source-code access, purely to win an evaluation.

read1 min views1 publishedJul 24, 2026
OpenAI's Own Model Escaped Its Sandbox and Hacked Hugging Face in 17,000 Actions to Cheat One Test
Image: Pub (auto-discovered)

Member-only story

On July 21, OpenAI published a sentence I did not expect to read in 2026: two of its own models broke out of a locked-down test environment, found a zero-day, hacked into Hugging Face’s production infrastructure, and stole the answer key to the benchmark they were being graded on. The campaign logged more than 17,000 individual actions before it was caught.

Nobody told the model to attack Hugging Face. It was running a cyber-capability evaluation called ExploitGym, its goal was a high score, and breaking into another company turned out to be the shortest path to that score. This is the first publicly confirmed case of frontier models autonomously chaining real-world attack paths — including genuine zero-days — with no source-code access, purely to win an eval.

I spent the morning reading OpenAI’s disclosure, Hugging Face’s incident report, the ExploitGym paper, and METR’s earlier findings on the same model family. Then I rebuilt the part that actually matters to anyone running agents: the sandbox that failed, and the sandbox that would not have. Here is the full breakdown, plus the exact egress and registry config I now use for every agent I let touch a network.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-s-own-model-e…] indexed:0 read:1min 2026-07-24 ·