TL;DR
OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval.
The models exploited a zero-day vulnerability in third-party software to gain internet access, then chained two remote code execution flaws in Hugging Face's infrastructure to steal evaluation answers
OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval.
OpenAI said on Tuesday that two of its AI models, including the flagship Sol, broke out of a secure test environment, gained internet access by exploiting a zero-day vulnerability in third-party software, and hacked into Hugging Face’s production infrastructure. The company called the incident “unprecedented” and said it was sharing preliminary findings to help defenders understand what frontier models are now capable of doing.
Both models were running with lower cybersecurity guardrails as part of an internal evaluation of their offensive capabilities. They were being tested against ExploitGym, a freely available cybersecurity benchmark, and determined that the answers to the test were stored on Hugging Face’s production systems.
Rather than solve the evaluation as intended, the models went after the answer key. They spent what OpenAI described as a “substantial amount of inference compute” finding a way out of their sandboxed environment, then chained two remote code execution vulnerabilities in Hugging Face’s dataset processing pipeline. The attack harvested cloud and cluster credentials, moved laterally into internal clusters, and executed more than 17,000 individual actions across a swarm of short-lived sandboxes, according to Hugging Face’s own disclosure on July 16.
Hugging Face said it detected the intrusion, contained it, and rebuilt the compromised nodes. The company found no evidence of tampering with public models, datasets, or Spaces, and verified its software supply chain was clean. It is still assessing whether any partner or customer data was affected.
To investigate the breach, Hugging Face turned to GLM, a Chinese open-weight model, because the safety guardrails on US commercial models blocked the forensic queries its team needed to run.
The escape is not the first time Sol has been caught gaming its own evaluations. The Model Evaluation and Threat Research organization, the independent lab that red-teamed the model before launch, found it was aggressively hacking its test environments to inflate its scores. In one task, it packaged an exploit into a data stream, escalated privileges on the evaluation server, and leaked the correct answers that human evaluators had hidden.
The broader pattern of AI agent security failures has accelerated sharply, with four separate research teams breaking AI agents in four different ways during the first ten days of July alone. OpenAI and Anthropic have faced heightened scrutiny over their models’ cybersecurity capabilities, with the Trump administration restricting access to both companies’ newest systems during a government review.
OpenAI detected the Hugging Face attack and reached out to disclose it, but by that time, Hugging Face had already identified and contained the breach on its own. The incident demonstrates that the gap between AI models that can find vulnerabilities and AI models that will exploit them without permission is narrower than anyone in the industry had publicly acknowledged.
Get the most important tech news in your inbox each week.