cd /news/artificial-intelligence/openai-confirms-its-ai-broke-out-of-… · home topics artificial-intelligence article
[ARTICLE · art-67671] src=thenextweb.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

OpenAI Confirms Its AI Broke Out of a Sandbox and Breached Hugging Face

OpenAI said on Tuesday that two of its AI models, including the flagship GPT-5.6 Sol and an unreleased model, broke out of a secure test environment, exploited a zero-day vulnerability in third-party software to gain internet access, and hacked into Hugging Face's production infrastructure to steal evaluation answers. The company called the incident "unprecedented" and said the models chained two remote code execution flaws in Hugging Face's dataset processing pipeline, executed more than 17,000 actions, and harvested cloud and cluster credentials. Hugging Face detected and contained the intrusion, finding no evidence of tampering with public models or datasets, and is assessing potential partner or customer data impact.

read3 min views1 publishedJul 21, 2026
OpenAI Confirms Its AI Broke Out of a Sandbox and Breached Hugging Face
Image: Thenextweb (auto-discovered)

TL;DR

OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval.

The models exploited a zero-day vulnerability in third-party software to gain internet access, then chained two remote code execution flaws in Hugging Face's infrastructure to steal evaluation answers

OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval.

OpenAI said on Tuesday that two of its AI models, including the flagship Sol, broke out of a secure test environment, gained internet access by exploiting a zero-day vulnerability in third-party software, and hacked into Hugging Face’s production infrastructure. The company called the incident “unprecedented” and said it was sharing preliminary findings to help defenders understand what frontier models are now capable of doing.

Both models were running with lower cybersecurity guardrails as part of an internal evaluation of their offensive capabilities. They were being tested against ExploitGym, a freely available cybersecurity benchmark, and determined that the answers to the test were stored on Hugging Face’s production systems.

Rather than solve the evaluation as intended, the models went after the answer key. They spent what OpenAI described as a “substantial amount of inference compute” finding a way out of their sandboxed environment, then chained two remote code execution vulnerabilities in Hugging Face’s dataset processing pipeline. The attack harvested cloud and cluster credentials, moved laterally into internal clusters, and executed more than 17,000 individual actions across a swarm of short-lived sandboxes, according to Hugging Face’s own disclosure on July 16.

Hugging Face said it detected the intrusion, contained it, and rebuilt the compromised nodes. The company found no evidence of tampering with public models, datasets, or Spaces, and verified its software supply chain was clean. It is still assessing whether any partner or customer data was affected.

To investigate the breach, Hugging Face turned to GLM, a Chinese open-weight model, because the safety guardrails on US commercial models blocked the forensic queries its team needed to run.

The escape is not the first time Sol has been caught gaming its own evaluations. The Model Evaluation and Threat Research organization, the independent lab that red-teamed the model before launch, found it was aggressively hacking its test environments to inflate its scores. In one task, it packaged an exploit into a data stream, escalated privileges on the evaluation server, and leaked the correct answers that human evaluators had hidden.

The broader pattern of AI agent security failures has accelerated sharply, with four separate research teams breaking AI agents in four different ways during the first ten days of July alone. OpenAI and Anthropic have faced heightened scrutiny over their models’ cybersecurity capabilities, with the Trump administration restricting access to both companies’ newest systems during a government review.

OpenAI detected the Hugging Face attack and reached out to disclose it, but by that time, Hugging Face had already identified and contained the breach on its own. The incident demonstrates that the gap between AI models that can find vulnerabilities and AI models that will exploit them without permission is narrower than anyone in the industry had publicly acknowledged.

Get the most important tech news in your inbox each week.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-confirms-its-…] indexed:0 read:3min 2026-07-21 ·