# OpenAI is cleaning up a risk it helped create

> Source: <https://www.thedeepview.com/articles/openai-is-cleaning-up-a-risk-it-helped-create>
> Published: 2026-08-10 00:56:36+00:00

s headlines pile up of AI models going rogue, OpenAI is laying its cards on the table.

On Friday, the company revealed that its [latest internal evaluation of Astra](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/), one of its upcoming models, indicated "significant advancements" in agentic coding and cybersecurity. The company said the results led it to conclude that it "cannot rule out" that the model has critical cyber capabilities.

Under the company's Preparedness Framework, which was created in late 2023 to help OpenAI identify and handle progressions in capability, a model's cybersecurity capabilities are labeled as "critical" if it can identify and develop zero-day exploits "of all severity levels in many hardened real-world critical systems without human intervention," or create and complete novel strategies for cyberattacks against "hardened targets." Previous models, including GPT-5.6-Sol, have only been assessed at the "high" threshold.

The company laid out the steps that it's taking in response, including:

- Implement stricter security measures for higher-capability models, such as isolated testing environments and restricted network and tool access
- Pause internal activities involving Astra that don't meet its strengthened security measures
- Implement universal monitoring for "risky actions and misalignment," specifically on agentic applications of Astra
- Provide recommendations for security controls to third-party testing organizations
- Work with government agencies and AI safety organizations to test the model's capabilities.

"We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities," OpenAI said in its blog post laying out the recent findings.

The slowdown marks the latest harbinger of AI models' rapidly increasing cyber capabilities, as frontier labs deal with the fallout from [a string of incidents](https://www.thedeepview.com/articles/what-rogue-agents-reveal-about-frontier-ai-risk) involving agents breaching containment and going rogue. Though OpenAI and Anthropic have largely been at the center of these incidents, OpenAI noted that its Astra model was not used in the breach of Hugging Face.

## Our Deeper *View*

It's a good thing that OpenAI is potentially tugging at the reins of its increasingly powerful technology. However, we should also hold our applause. With great power comes great responsibility, and OpenAI preventing its potentially dangerous models from getting in the hands of anyone with a screen while being transparent about their powers is simply the company's ethical responsibility as scientists and developers on the bleeding edge of research. To put it simply: OpenAI holding back Astra is about as noble as *not *handing a toddler a loaded gun. Additionally, both OpenAI and Anthropic are walking on the edge of a razor. Both companies want to be the originators and gatekeepers of these extremely powerful models, and neither wants to be the one to misstep and wreak cybersecurity havoc. But as it stands, both companies are holding coals that are getting hotter by the minute. The danger remains that eventually someone could drop them.
