{"slug": "openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards", "title": "OpenAI says Astra could reach ‘critical’ cyber capability, tightens safeguards", "summary": "OpenAI said its upcoming model Astra could reach 'critical' cyber capability, the highest risk category in its Preparedness Framework, where a system can autonomously find and exploit vulnerabilities or carry out end-to-end cyberattacks against hardened targets. The company, citing internal testing and expert reviews, said it cannot rule out critical cyber capabilities and is tightening security controls, including isolated testing environments and enhanced monitoring. Analysts warn that enterprise security must evolve from reactive to preemptive as AI-driven attacks become more feasible.", "body_md": "OpenAI said its upcoming model Astra is showing cybersecurity capabilities that could reach its highest risk category, where a system can autonomously find and exploit vulnerabilities or carry out end-to-end cyberattacks against hardened targets.\n\nThe company disclosed the assessment following recent internal testing and expert reviews.\n\n“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI said in a [statement](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/). “These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.\n\nTo explain the shift, OpenAI pointed to its internal [Preparedness Framework](https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf), which tracks how far AI models advance in sensitive areas such as cybersecurity.\n\nAt the top end of that framework are systems that no longer just assist humans but can act on their own, the company said.\n\n“A model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal,” the statement added.\n\nThe company said Astra has not yet been definitively classified at that level, but its early performance is “strong enough” that such a designation cannot be ruled out. Its earlier models, including GPT 5.6 Sol, “have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold.”\n\n“This is a substantial inflection point,” said Apeksha Kaushik, senior principal analyst at Gartner. “An AI system could autonomously discover vulnerabilities, develop exploits, and execute end-to-end attacks with minimal human guidance.”\n\nFor security teams, the change is not just technical — it affects how attacks may unfold, analysts feel.\n\nKaushik said the pace of progress suggests “practical, real-world exploitation is becoming increasingly feasible,” meaning attackers could automate large parts of the attack process. That reduces the time defenders have to react.\n\n“The implication is clear: enterprise security must evolve from reactive to preemptive,” she said, adding that organizations should move toward continuous, AI-driven exposure assessment and predictive analysis.\n\nSanchit Vir Gogia, chief analyst at Greyhound Research, said companies should not wait for a formal label before acting.\n\n“OpenAI has said it cannot rule out critical cybersecurity capability in Astra and is treating the model accordingly. That is a precautionary trigger rather than a finished finding,” he said.\n\nHe added that the focus should shift beyond patching speed. “The measure that matters is defensive response latency. A flat vulnerability queue is no longer a security posture.”\n\nBased on the development, OpenAI said it is tightening controls around Astra’s development.\n\n“We are implementing stricter security controls for higher-capability models,” the company said, including “isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution,” OpenAI added I n the statement.\n\nIt is also “pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.”\n\nThe company said it has expanded monitoring across how the model is used. “We have implemented universal monitoring for risky actions and misalignment,” it said, adding that systems can “trigger a security response to review and interrupt high-risk activity.”\n\nAnalysts said these steps are necessary, but may not fully address the risks as capabilities improve.\n\n“Current safeguards such as restricted environments, continuous monitoring, and external red-teaming are necessary, but the gap between safeguards and emerging threats is increasing,” Kaushik said. She pointed to risks such as prompt injection and weak access controls in AI systems.\n\nGogia said safeguards need to be viewed in the context of the broader system.\n\n“A capable model does not operate inside a framework document. It operates inside a system, and systems leak authority through their exceptions,” he said. “Gated access buys defenders time. It does not repeal a capability.”\n\nOpenAI said it will work with governments and external safety groups to further test Astra.\n\n“We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model,” the company said, adding that it will also share guidance with third-party testing partners. The company said it is disclosing the findings to be transparent about what it called a “potential shift in capabilities.”", "url": "https://wpnews.pro/news/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards", "canonical_source": "https://www.csoonline.com/article/4207311/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards.html", "published_at": "2026-08-10 12:07:10+00:00", "updated_at": "2026-08-10 12:09:04.607621+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-policy"], "entities": ["OpenAI", "Astra", "GPT 5.6 Sol", "Apeksha Kaushik", "Gartner", "Sanchit Vir Gogia", "Greyhound Research"], "alternates": {"html": "https://wpnews.pro/news/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards", "markdown": "https://wpnews.pro/news/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards.md", "text": "https://wpnews.pro/news/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards.txt", "jsonld": "https://wpnews.pro/news/openai-says-astra-could-reach-critical-cyber-capability-tightens-safeguards.jsonld"}}