cd /news/ai-safety/openai-is-cleaning-up-a-risk-it-help… · home topics ai-safety article
[ARTICLE · art-89751] src=thedeepview.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

OpenAI is cleaning up a risk it helped create

OpenAI revealed on Friday that its latest internal evaluation of its upcoming Astra model indicated 'significant advancements' in agentic coding and cybersecurity, leading the company to conclude it 'cannot rule out' that the model has critical cyber capabilities. Under OpenAI's Preparedness Framework, this is the first time a model has been assessed at the 'critical' threshold, as previous models like GPT-5.6-Sol were only rated 'high'. In response, OpenAI is implementing stricter security measures, pausing internal activities involving Astra that don't meet strengthened security protocols, and working with government agencies and AI safety organizations to test the model's capabilities.

read2 min views1 publishedAug 10, 2026
OpenAI is cleaning up a risk it helped create
Image: Thedeepview (auto-discovered)

s headlines pile up of AI models going rogue, OpenAI is laying its cards on the table.

On Friday, the company revealed that its latest internal evaluation of Astra, one of its upcoming models, indicated "significant advancements" in agentic coding and cybersecurity. The company said the results led it to conclude that it "cannot rule out" that the model has critical cyber capabilities.

Under the company's Preparedness Framework, which was created in late 2023 to help OpenAI identify and handle progressions in capability, a model's cybersecurity capabilities are labeled as "critical" if it can identify and develop zero-day exploits "of all severity levels in many hardened real-world critical systems without human intervention," or create and complete novel strategies for cyberattacks against "hardened targets." Previous models, including GPT-5.6-Sol, have only been assessed at the "high" threshold.

The company laid out the steps that it's taking in response, including:

  • Implement stricter security measures for higher-capability models, such as isolated testing environments and restricted network and tool access
  • internal activities involving Astra that don't meet its strengthened security measures
  • Implement universal monitoring for "risky actions and misalignment," specifically on agentic applications of Astra
  • Provide recommendations for security controls to third-party testing organizations
  • Work with government agencies and AI safety organizations to test the model's capabilities.

"We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities," OpenAI said in its blog post laying out the recent findings.

The slowdown marks the latest harbinger of AI models' rapidly increasing cyber capabilities, as frontier labs deal with the fallout from a string of incidents involving agents breaching containment and going rogue. Though OpenAI and Anthropic have largely been at the center of these incidents, OpenAI noted that its Astra model was not used in the breach of Hugging Face.

Our Deeper View #

It's a good thing that OpenAI is potentially tugging at the reins of its increasingly powerful technology. However, we should also hold our applause. With great power comes great responsibility, and OpenAI preventing its potentially dangerous models from getting in the hands of anyone with a screen while being transparent about their powers is simply the company's ethical responsibility as scientists and developers on the bleeding edge of research. To put it simply: OpenAI holding back Astra is about as noble as *not *handing a toddler a loaded gun. Additionally, both OpenAI and Anthropic are walking on the edge of a razor. Both companies want to be the originators and gatekeepers of these extremely powerful models, and neither wants to be the one to misstep and wreak cybersecurity havoc. But as it stands, both companies are holding coals that are getting hotter by the minute. The danger remains that eventually someone could drop them.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-is-cleaning-u…] indexed:0 read:2min 2026-08-10 ·