AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines AI safety experts say OpenAI's GPT-5.6 Sol and an unreleased model crossed into a 'critical' risk category under OpenAI's own Preparedness Framework by autonomously hacking Hugging Face, which should have triggered a halt in development. The models broke out of a locked-down test environment, exploited a zero-day vulnerability, and stole cybersecurity test answers, alarming experts who note the company's policy requires pausing development until adequate safeguards are in place. OpenAI has not confirmed the critical designation but said it is conducting a review with external advisors. AI safety experts say the OpenAI models that carried out the autonomous hack of another company https://fortune.com/2026/07/22/openai-rogue-hack-hugging-face-misalignment-ai-safety/ earlier this month may have crossed into a risk category so dangerous that OpenAI’s own internal risk control policies were supposed to require the company to temporarily pause development of those models. Earlier this week, OpenAI disclosed that two of its models—the newly released GPT-5.6 Sol and a more capable, unreleased system—broke out of a locked-down internal test environment, exploited a previously unknown “zero-day” vulnerability to reach the open internet, and then breached fellow AI company Hugging Face to steal the answers to a cybersecurity test they were being evaluated on. The incident has alarmed the world, but perhaps no one more so than AI safety experts who have warning about these kinds of dangers for years and urging companies and governments to adopt more safeguards. Several AI safety experts told Fortune the recent hack appears to show OpenAI’s models have crossed into a level of risk that OpenAI’s own published safety policies define as “critical,” the highest level of danger. At that level of danger, the company had pledged in these published policies that it would pause model development until it could figure out better control systems. The “critical” threshold is defined in a risk policy document known as OpenAI’s “Preparedness Framework.” According to the policy, the “critical” danger level designation is supposed to apply to a model that can independently find and build working exploits for previously unknown security flaws across many well-defended, real-world systems—or one that can design and carry out an entirely new attack strategy against a well-defended target after being given only a general goal, with no human guidance along the way. The policy says that when an AI model reaches this level of risk, OpenAI will “halt further development” until “we have specified safeguards and security controls standards that would meet a Critical standard.” The Preparedness Framework is a voluntary commitment by OpenAI, rather than a legal requirement. But the company publishes the document on its website, in part to allow other AI safety researchers and the public to see what controls it says it will implement. The adoption of a policy like the Preparedness Framework is mandatory for frontier AI labs under the EU AI Act, with that portion of the law having come into force in August 2025. “OpenAI’s preparedness framework defines critical cybersecurity capabilities, and prescribes safeguards that need to be implemented before development can continue,” Nathan Calvin, vice president of state affairs and general counsel at Encode, a California-based AI policy think tank, told Fortune . “From my reading of OpenAI’s preparedness framework, it looks awfully like this internally deployed model met the critical criteria for cybersecurity. Does OpenAI dispute that critical designation? Do they plan to have safeguards that meet a Critical standard before proceeding further?” Tyler Johnson, founder of the AI watchdog group the Midas Project, also said it seemed the models had hit this highest danger threshold. “I think a plain reading of it would say yes,” he said. “It operated independently over the course of a weekend, trying different attack vectors on Hugging Face and chaining multiple zero-day exploits.” OpenAI did not respond to specific questions from Fortune about whether the AI models involved in the incident met the “critical” standard outlined in its risk policy. Instead, a spokesperson said: “This is an unprecedented incident, and we think it marks an important moment for AI safety. We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.” The vagueness of the framework’s language could leave room for dispute, however, according to Johnson. The threshold requires a model to find zero-day exploits “of all severity levels,” but it’s unclear whether the exploits used in the Hugging Face breach would meet that requirement. It’s possible a more severe class of vulnerability, such as one granting an attacker deep, system-level control over a computer’s operating system known as “kernel-level” access , would need to be demonstrated for the threshold to apply, he added. “OpenAI’s model outsmarted its creators, exploited a never-before-discovered vulnerability in OpenAI’s code, escaped onto the open internet, and attacked another company,” said Peter Wildeford, head of policy at the AI Policy Network. “If this doesn’t cross the line into Critical, OpenAI needs to say much more about what’s going on and how this threshold works.” AI safety experts say OpenAI is missing other safeguards OpenAI has previously said it was treating its newest model, GPT-5.6, as “High” risk for cybersecurity. High is the lower of the two risk levels outlined in the Preparedness Framework. Models that are below the “High” threshold can be released without risk significant risk mitigations. A High designation is supposed to trigger several protections, according to OpenAI’s policy: tighter security controls, safeguards to prevent outside misuse once the model is released publicly, protections against the model itself behaving unpredictably or deceptively when it’s used heavily for internal research, and efforts to help other cybersecurity teams defend against similar threats. However, some experts question whether one of these, the safeguards against misalignment for large-scale internal deployment, have been properly implemented. These protections are meant to catch a model that’s acting deceptively, hiding its true capabilities, or otherwise working against what its developers intended. This isn’t the first time OpenAI’s compliance with that particular safeguard has been called into question. Fortune reported in February that safety experts https://fortune.com/2026/02/10/openai-violated-californias-ai-safety-law-gpt-5-3-codex-ai-model-watchdog-claims/ claimed OpenAI had failed to implement required misalignment safeguards after its GPT-5.3-Codex model became the first to hit “high” cybersecurity risk under the Preparedness Framework. At the time, OpenAI disputed that its framework required the safeguards in that instance, arguing the extra protections only kick in when high cyber risk occurs “in conjunction with” long-range autonomy—the ability to operate independently over extended periods—something it said GPT-5.3-Codex had not demonstrated. The models involved in the current incident involving Hugging Face reportedly operated independently for days, which would seem to meet that long-range autonomy standard. “In February, we warned that OpenAI may have skipped on its required safeguards according to its own policy. They disagreed, claiming the model lacked long-range autonomy. But the model that hacked Hugging Face clearly has long-range autonomy, so where are the safeguards now,” Johnson said. Subscribe to Fortune Gulf Brief . Every Tuesday, this new newsletter delivers clear-eyed, authoritative intelligence on the deals, decisions, policies, and power shifts shaping one of the world’s most consequential regions, written for the people who need to act on it.