# OpenAI has already ended an internal pause

> Source: <https://www.lesswrong.com/posts/k3eKqKzq4Y7xnqEfZ/openai-has-already-ended-an-internal-pause>
> Published: 2026-07-31 12:03:18+00:00

One day before OpenAI’s HF incident disclosure, OpenAI[ disclosed](https://openai.com/index/safety-alignment-long-horizon-models/) that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a resumption decision has already been made against a standard that has not really been formalized. We need to prevent this from happening again.

*Epistemic status: could have been a short-form.*

[OpenAI, 20th July](https://openai.com/index/safety-alignment-long-horizon-models/): *"To evaluate the new monitoring system, we replayed a small set of internal deployment environments where the model previously pursued misaligned actions, this time with the new safeguards in place. The new safeguards were able to catch considerably more misaligned actions pursued by the model, and the ones it missed were all judged to be low-severity."*

*"After testing the new system, we concluded that limited internal access to models with long-horizon capabilities could be restored. We have not observed any serious circumvention of safeguards since redeployment began several weeks ago. The first version of these safeguards was deliberately conservative. We have continued tuning the system to reduce unnecessary interruptions without weakening the safeguards."*

…

One day later, OpenAI [announced a bold partnership](https://openai.com/index/hugging-face-model-evaluation-security-incident/) with Hugging Face.

From that post: "*These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.*"

The safeguards self-certified as adequate on the 20th were off during the evaluation described on the 21st.

Their framework says a *critical cyber determination means halting development*.

Here's the exit condition: "*until we have specified safeguards and security controls that would meet a Critical standard*"

This is completely circular.

The review goes to their Safety and Security Committee. Fair enough. But the process can be followed perfectly, and we still learn nothing, because the standard it applies has never been published.

**For frontier companies: Publish the criteria before the determination, not after.** OpenAI, this is what you said you would be doing when approaching those levels of capability.

**For Lesswrong folks: we should be debating what those standards are now**, otherwise we get mitigations that hold for a year, then fail against a much more capable model. There is basically no literature on the matter. [1] Otherwise, everything will be done in an ad-hoc way.

…what an adequate safeguard should be. Across the developer frameworks,[ METR's policy comparison](https://metr.org/common-elements), GovAI's safety-case work, RAND and the[ GPAI Code of Practice](https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai), every pre-committed number is basically describing the level of capability needed to force action. On the response side, after running a long literature review with Claude, I was not able to find numbers, or even solid framework stating what would have to be true to resume.

The recurring phrasing is "reduced to an *acceptable* level", acceptable being again tautological. The one real exception is security-only:[ RAND's SL1-SL5](https://www.rand.org/pubs/research_reports/RRA2849-1.html) for weight protection, which Anthropic and Google DeepMind both map to. The closest proposal is[ Alaga and Schuett's](https://arxiv.org/abs/2310.00374) two-threshold scheme, which puts the resumption bar above the trigger bar but leaves adequacy undefined per capability. The[ Frontier Model Forum](https://www.frontiermodelforum.org/technical-reports/risk-taxonomy-and-thresholds/) reaches the same structure and says so: "more research is needed to determine… what kind of safety precautions are *adequate*."
