{"slug": "openai-admits-monitorability-has-tumbled-for-its-new-astra-model", "title": "OpenAI admits \"monitorability\" has tumbled for its new Astra model", "summary": "OpenAI has begun rolling out its GPT-6 Astra model to selected customers, warning that the model shows a substantial decrease in chain-of-thought monitorability compared to previous models, making it harder to detect misbehavior. The UK's AI Safety Institute independently confirmed the increased opacity, finding that Astra can solve harder problems without showing reasoning and that up to 80% of reasoning summaries were missing on long simulated cyber trajectories. OpenAI chief scientist Jakub Pachocki expressed concern, stating that no lab has solved alignment and monitoring sufficiently to continue responsibly scaling at maximum speed.", "body_md": "OpenAI has started rolling out its most recent Astra model to [selected customers](https://openai.com/index/gpt-6-astra/?ref=thestack.technology) today with a warning: it’s a lot harder to monitor reasoning. \n\n“According to our evaluations, GPT-6 Astra shows a substantial decrease in chain-of-thought [CoT] monitorability compared to previous models,” the [GPT-6 Astra model card said.](https://deploymentsafety.openai.com/gpt-6-astra/aggregate-monitorability-findings?ref=thestack.technology) \n\nThe model has \"significantly lower full-context monitorability across most output token lengths\" than GPT 5.6 Sol, it added\n\nBeing able to understand and observe model behaviour has become increasingly important as models act with more autonomy.\n\nIt is also critical under some regulatory regimes.\n\n\"I am deeply worried by the trend of decreasing CoT monitorability: CoT monitoring is a core part of our misalignment safety strategy that has no good substitute now,\" OpenAI technical staff member Tomek Korbak posted on X.\n\nHe added: \"We will continue closely tracking CoT monitorability, investigating the reasons why it’s decreasing and trying to reverse the trend.\"\n\nThe Astra model card also states that the model, when told chain-of-thought monitoring is present or explicitly directed to evade monitoring, can control its chain-of-thought reasoning traces to decrease detection.\n\nOpenAI bluntly states the shift in chain-of-thought monitorability represents \"a decrease in the extent to which Astra’s chain of thought contains incriminating information about the model’s misbehavior.\"\n\nThe model often produces \"shorter CoTs that often omit or weaken the evidence the monitor needs, including by producing empty or nearly empty CoTs more often,\" it said.\n\nOpenAI said it was confident that architectural changes in the model hadn't caused the shift. The increase correlates with the model's ability to perform tasks without \"verbalising\" them first; i.e. with non-CoT outputs.\n\nThe increased opacity of Astra prompted OpenAI’s chief scientist Jakub Pachocki to publish an essay [__on Sunday__](https://openai.com/index/an-alien-mind/?ref=thestack.technology) via the company blog warning about \"unchecked alien intelligence\" and calling for other AI labs to work together on improving alignment and monitoring.  \n\n“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” Pachocki wrote.\n\nSome 24 hours later, his employer announced it would be rolling out the new model to selected customers and all paying subscribers in the “next few days.”\n\n### **AISI confirms opaqueness** \n\nThe Astra model card includes an independent review from Britain's AI Safety Institute that confirm OpenAI’s presentation of the model’s monitoring issues.\n\nThe institute found that the model was more opaque than its predecessors when it came to the ability of users to monitor its reasoning.\n\nThe UK’s AI minister Kanishka Narayan commented [__on the findings in a post on LinkedIn__](https://lnkd.in/p/eQn4HnnQ?ref=thestack.technology) on Saturday. “AISI’s monitorability testing found that Astra had capabilities that could help it avoid detection by monitoring systems. In particular, Astra can solve significantly harder problems without showing its reasoning and has greater control over what appears in that reasoning,” he wrote.\n\nNarayan said the independent review also found Astra’s raw reasoning was \"more compressed and sometimes harder to interpret.\"\n\nThe model card reads, “During AISI’s evaluations, reasoning summaries were not consistently provided by the user API, with up to 80% missing on long simulated cyber trajectories.\"\n\nNarayan caveated it was too early to say whether these findings represented a trend in future models, but added it was important to be vigilant for potential harms.\n\n### **Race to the bottom**\n\nOpenAI’s RSI Preparedness lead Micah Carroll said [on X](https://x.com/MicahCarroll/status/2095603855316996529?ref=thestack.technology) after the model’s announcement last week: “GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation.” \n\nCarroll made the same point at Pachocki highlighting it is in “everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom.\"\n\n\"Nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected,” Carroll wrote.\n\nThis has not stopped OpenAI from releasing the model to a \"limited set of organizations\" today and rolling it out to all paying subscribers “in the coming days,” according to its blog announcement published on Monday.\n\nThe degraded monitorability may be a sticking point for EU organisations wanting to test out the new Astra model.\n\nThe [__EU AI Act__](https://artificialintelligenceact.eu/high-level-summary/?ref=thestack.technology), which was updated on August 31, prohibits AI systems that deploy “subliminal, manipulative, or deceptive techniques” that can result in causing harm or “impairing informed decision-making.”\n\nProviders of high-risk AI systems, i.e systems that are a safety component or a regulated product under EU law, must design them with necessary logging requirements, so any risks to health, safety or fundamental rights caused by the system can be monitored, per the act.\n\nOpenAI said it was committed to tracking this degradation in monitorability and would “not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization.”", "url": "https://wpnews.pro/news/openai-admits-monitorability-has-tumbled-for-its-new-astra-model", "canonical_source": "https://www.thestack.technology/open-ai-astra-monitor-warning/", "published_at": "2026-09-07 12:45:57+00:00", "updated_at": "2026-09-07 12:58:01.065261+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "large-language-models", "ai-research"], "entities": ["OpenAI", "GPT-6 Astra", "GPT 5.6 Sol", "Tomek Korbak", "Jakub Pachocki", "AI Safety Institute", "Kanishka Narayan", "Micah Carroll"], "alternates": {"html": "https://wpnews.pro/news/openai-admits-monitorability-has-tumbled-for-its-new-astra-model", "markdown": "https://wpnews.pro/news/openai-admits-monitorability-has-tumbled-for-its-new-astra-model.md", "text": "https://wpnews.pro/news/openai-admits-monitorability-has-tumbled-for-its-new-astra-model.txt", "jsonld": "https://wpnews.pro/news/openai-admits-monitorability-has-tumbled-for-its-new-astra-model.jsonld"}}