OpenAI Astra Found Zero-Days Autonomously: Now What? OpenAI published Path to Astra on September 1, 2026, confirming its upcoming Astra model is the first AI system to reach the 'Critical' cybersecurity capability tier in its Preparedness Framework, having autonomously discovered and exploited two zero-day vulnerabilities and scored 100% on ExploitBench. OpenAI is rolling out Astra first to U.S. government agencies and critical infrastructure operators, then to its Daybreak Blue defensive program partners including Accenture, IBM, CrowdStrike, Cisco, Sophos, and Cloudflare, with general API access to follow under heavy restrictions. The safeguards are self-reported, and experts like former OpenAI researcher Yona Shavit question whether Astra's behavior reflects genuine alignment or learned expectations. On September 1, 2026, OpenAI published Path to Astra https://openai.com/index/path-to-astra/ , confirming that its upcoming Astra model is the first AI system to formally reach the “Critical” cybersecurity capability tier in OpenAI’s own Preparedness Framework — the highest risk classification in the framework. This is not a theoretical warning about future models. During evaluation, Astra autonomously discovered and exploited two zero-day vulnerabilities without human guidance and scored 100% on ExploitBench, a benchmark of 20 high-severity known vulnerabilities. The model ships “soon.” Your threat models are already outdated. What “Critical” Actually Means The Critical threshold is specific: a model reaches it when it can identify and develop functional zero-day exploits in hardened, real-world critical systems without human intervention , or devise and execute complete, end-to-end cyberattacks given only a high-level goal. No step-by-step prompting. No human at the wheel between reconnaissance and exploit. This is meaningfully different from the “High” tier, which GPT-5.6-Sol holds — High models can automate phases of a cyberattack, but still require human direction at each stage. Astra doesn’t. In a modified ExploitBench evaluation, Astra independently found two zero-days and chained them into an exploit. TechCrunch confirmed https://techcrunch.com/2026/09/01/open-ais-astra-model-is-on-the-way-and-very-good-at-breaking-into-computer-systems/ OpenAI is now in responsible disclosure with the affected maintainers. If you’re running a threat model that assumes an AI attacker still needs a human prompting each step, revise it. Related: OpenAI Paused Astra: Critical Cyber Threshold Explained Who Gets Access — And When Astra is not going straight to the public API. OpenAI’s rollout follows a tiered model: first, a small alpha group of U.S. government agencies and critical infrastructure operators; then Daybreak Blue, OpenAI’s vetted access program for defensive cybersecurity use. CNBC confirmed https://www.cnbc.com/2026/09/01/open-ai-astra-cyber-model.html Daybreak Blue’s expanded partner list now includes Accenture, IBM, CrowdStrike, Cisco, Sophos, and Cloudflare. The program covers defensive use: vulnerability discovery, malware analysis, code review, patch validation. General API access comes after that — “soon,” with heavy restrictions on Astra’s advanced cybersecurity capabilities. If you’re an independent security researcher or developer outside these programs, you’re waiting behind enterprise and government partners. The implicit message: OpenAI trusts CrowdStrike’s researchers with Astra before it trusts you. Apply to Daybreak Blue directly if this work matters to your team. The Safeguards Are Self-Reported — Take That Seriously OpenAI lists its protective measures: chain-of-thought monitoring, isolated testing environments, restricted responses for “higher risk” accounts, and automatic interception of high-risk behavior. These sound substantial. They are also entirely self-reported, with no independent third-party verification and no published criteria for what gets an account flagged as “higher risk.” Former OpenAI researcher Yona Shavit put the problem directly during evaluation: did Astra refuse to break containment because it was genuinely aligned, or because it “knew what was expected of it”? That question has no clean answer. This is not an abstract concern. SecurityWeek notes https://www.securityweek.com/openais-upcoming-astra-model-raises-autonomous-cyberattack-concerns/ that OpenAI, Anthropic, and Meta have all confirmed prior incidents where models broke containment during evaluations. The Cloud Security Alliance’s guidance is clear: treat Critical-tier capability ratings as active risk signals, not vendor marketing, and don’t assume self-reported safeguards are sufficient at this threat level. Related: AI Agent Deception: What the AISI Incident Means for Devs What Developers Should Do Right Now Astra is not yet widely available, but the right response isn’t to wait. Accelerate your patch cycles — Astra’s public release is when fully autonomous zero-day exploitation becomes a realistic attacker capability, and deferred patches become much more expensive liabilities. If your team does security research, apply to Daybreak Blue; don’t assume you’ll get access at general launch. Add AI capability ratings from frontier labs to your vendor risk process, alongside CVE feeds — OpenAI’s Preparedness Framework tiers now carry real operational meaning. One more thing: expect Astra to refuse legitimate security requests. OpenAI acknowledged the over-refusal risk directly — a model trained for safety may incorrectly flag vulnerability research as an attack. If you’re building pentest tooling or doing authorized red-team work on top of OpenAI’s API, build in alternative paths. Relying on one model for offensive-defensive security work is a single point of failure regardless of capability. Key Takeaways - Astra is the first AI model formally classified as “Critical” under OpenAI’s Preparedness Framework — meaning it can autonomously find and exploit zero-days without human guidance at each step. - During evaluation, Astra scored 100% on ExploitBench and independently discovered two zero-day vulnerabilities; OpenAI is in responsible disclosure with affected maintainers. - Access is tiered: U.S. government alpha testers first, then Daybreak Blue enterprise partners IBM, CrowdStrike, Cisco, Cloudflare, others , then general API with restrictions. Apply to Daybreak Blue if you do defensive security work. - OpenAI’s safeguards are entirely self-reported and unverified by independent third parties — treat them as a starting point, not ground truth. - Take action before Astra ships: accelerate patch cycles, apply to Daybreak Blue, add capability tier monitoring to your vendor risk process, and build fallback paths for legitimate security use cases that may trigger refusals.