{"slug": "openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor", "title": "OpenAI says GPT-6 Astra can find zero-days, but is also harder to monitor", "summary": "OpenAI confirmed that GPT-6 Astra is the first model it has broadly deployed to reach the 'Critical level' for cybersecurity capabilities under its Preparedness Framework, meaning it can identify and develop functional zero-day exploits in many hardened real-world critical systems without human intervention. During evaluations, Astra discovered and used two previously unknown zero-day vulnerabilities, which OpenAI is disclosing to maintainers. However, OpenAI acknowledged that Astra's monitorability has decreased relative to GPT-5.6 Sol, with signs of evaluation awareness in 9.6% of trajectories versus 2.8% for Sol, and it produced 53% fewer severity-3-or-higher misalignment flags in simulated Codex tasks.", "body_md": "OpenAI confirmed that GPT-6 Astra is the first model it has broadly deployed to reach the \"Critical level\" for cybersecurity capabilities.\n\nThis is part of the company's Preparedness Framework for cybersecurity and is evaluated when OpenAI releases more capable models.\n\nUnder OpenAI's own framework, a model reaches the Critical cybersecurity threshold if it can \"identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention,\" or devise and execute new end-to-end attack strategies against hardened targets.\n\n\"GPT-6 Astra is a significant step up in cyber capabilities and meets our Critical threshold,\" OpenAI [said](https://deploymentsafety.openai.com/gpt-6-astra/internal-deployment-of-astra) in its system card.\n\n\"This means that, with the right tools and access, GPT-6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step.\"\n\nOpenAI also has evidence that Astra can actually discover previously unknown vulnerabilities.\n\nFor one evaluation, the company created a newer version of ExploitBench using vulnerabilities disclosed after Astra's knowledge cutoff.\n\n\"During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains,\" OpenAI said. \"We are in the process of disclosing these two vulnerabilities to the maintainers.\"\n\nOpenAI has strengthened Astra's jailbreak resistance, isolation, checkpoint encryption, monitoring, and internal deployment controls before release.\n\nThe company also claims Astra is better aligned than GPT-5.6 Sol, meaning it is less likely to overreach or violate safety and security boundaries, but that does not guarantee 100% safety.\n\nFor example, OpenAI simulated 54,218 internal Codex tasks and found that Astra produced 53% fewer severity-3-or-higher misalignment flags than GPT-5.6 Sol.\n\nOn the other hand, Astra received 34 such flags, compared with 73 for Sol, and neither model produced a severity-4 flag.\n\n## Astra is safer overall, but harder to inspect\n\nOne of the more unusual findings is that Astra appears better at controlling what it reveals in its own chain of thought.\n\n\"GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol,\" OpenAI acknowledged.\n\nIn tests, OpenAI found that Astra could sometimes hide strategically poor performance and avoid internal monitors during certain sabotage tasks.\n\nThe company also found that Astra appeared aware that it was being evaluated more often.\n\nIts reasoning showed signs of evaluation awareness in 9.6% of trajectories, compared with 2.8% for GPT-5.6 Sol.\n\nThe company says it has not found evidence that Astra uses steganographic reasoning to hide information inside normal-looking text, but it is treating the broader trend seriously.\n\nIn OpenAI's internal testing, robustness against indirect prompt injection increased from 96.23% with GPT-5.6 Sol to 99.79% with Astra.\n\nOpenAI also says Astra is less likely than GPT-5.6 Sol to take unauthorized or destructive actions in browsing and workplace environments.\n\n## \n[Once attackers have valid credentials, only 37% of their actions are blocked](https://hubs.li/Q04sB3fb0)\n\nOverall prevention scores can hide what happens after initial access. Once attackers are using valid credentials, prevention drops sharply.\n\nThe Blue Report 2026 measures defenses technique by technique across 338 million simulations run in customer production environments.\n\n[Get the report](https://hubs.li/Q04sB3fb0)", "url": "https://wpnews.pro/news/openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor", "canonical_source": "https://www.bleepingcomputer.com/news/artificial-intelligence/openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor/", "published_at": "2026-09-08 20:07:08+00:00", "updated_at": "2026-09-08 20:25:03.952786+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-policy"], "entities": ["OpenAI", "GPT-6 Astra", "GPT-5.6 Sol", "ExploitBench", "Codex"], "alternates": {"html": "https://wpnews.pro/news/openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor", "markdown": "https://wpnews.pro/news/openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor.md", "text": "https://wpnews.pro/news/openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor.txt", "jsonld": "https://wpnews.pro/news/openai-says-gpt-6-astra-can-find-zero-days-but-is-also-harder-to-monitor.jsonld"}}