# OpenAI admits "monitorability" has tumbled for its new Astra model

> Source: <https://www.thestack.technology/open-ai-astra-monitor-warning/>
> Published: 2026-09-07 12:45:57+00:00

OpenAI has started rolling out its most recent Astra model to [selected customers](https://openai.com/index/gpt-6-astra/?ref=thestack.technology) today with a warning: it’s a lot harder to monitor reasoning. 

“According to our evaluations, GPT-6 Astra shows a substantial decrease in chain-of-thought [CoT] monitorability compared to previous models,” the [GPT-6 Astra model card said.](https://deploymentsafety.openai.com/gpt-6-astra/aggregate-monitorability-findings?ref=thestack.technology) 

The model has "significantly lower full-context monitorability across most output token lengths" than GPT 5.6 Sol, it added

Being able to understand and observe model behaviour has become increasingly important as models act with more autonomy.

It is also critical under some regulatory regimes.

"I am deeply worried by the trend of decreasing CoT monitorability: CoT monitoring is a core part of our misalignment safety strategy that has no good substitute now," OpenAI technical staff member Tomek Korbak posted on X.

He added: "We will continue closely tracking CoT monitorability, investigating the reasons why it’s decreasing and trying to reverse the trend."

The Astra model card also states that the model, when told chain-of-thought monitoring is present or explicitly directed to evade monitoring, can control its chain-of-thought reasoning traces to decrease detection.

OpenAI bluntly states the shift in chain-of-thought monitorability represents "a decrease in the extent to which Astra’s chain of thought contains incriminating information about the model’s misbehavior."

The model often produces "shorter CoTs that often omit or weaken the evidence the monitor needs, including by producing empty or nearly empty CoTs more often," it said.

OpenAI said it was confident that architectural changes in the model hadn't caused the shift. The increase correlates with the model's ability to perform tasks without "verbalising" them first; i.e. with non-CoT outputs.

The increased opacity of Astra prompted OpenAI’s chief scientist Jakub Pachocki to publish an essay [__on Sunday__](https://openai.com/index/an-alien-mind/?ref=thestack.technology) via the company blog warning about "unchecked alien intelligence" and calling for other AI labs to work together on improving alignment and monitoring.  

“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” Pachocki wrote.

Some 24 hours later, his employer announced it would be rolling out the new model to selected customers and all paying subscribers in the “next few days.”

### **AISI confirms opaqueness** 

The Astra model card includes an independent review from Britain's AI Safety Institute that confirm OpenAI’s presentation of the model’s monitoring issues.

The institute found that the model was more opaque than its predecessors when it came to the ability of users to monitor its reasoning.

The UK’s AI minister Kanishka Narayan commented [__on the findings in a post on LinkedIn__](https://lnkd.in/p/eQn4HnnQ?ref=thestack.technology) on Saturday. “AISI’s monitorability testing found that Astra had capabilities that could help it avoid detection by monitoring systems. In particular, Astra can solve significantly harder problems without showing its reasoning and has greater control over what appears in that reasoning,” he wrote.

Narayan said the independent review also found Astra’s raw reasoning was "more compressed and sometimes harder to interpret."

The model card reads, “During AISI’s evaluations, reasoning summaries were not consistently provided by the user API, with up to 80% missing on long simulated cyber trajectories."

Narayan caveated it was too early to say whether these findings represented a trend in future models, but added it was important to be vigilant for potential harms.

### **Race to the bottom**

OpenAI’s RSI Preparedness lead Micah Carroll said [on X](https://x.com/MicahCarroll/status/2095603855316996529?ref=thestack.technology) after the model’s announcement last week: “GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation.” 

Carroll made the same point at Pachocki highlighting it is in “everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom."

"Nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected,” Carroll wrote.

This has not stopped OpenAI from releasing the model to a "limited set of organizations" today and rolling it out to all paying subscribers “in the coming days,” according to its blog announcement published on Monday.

The degraded monitorability may be a sticking point for EU organisations wanting to test out the new Astra model.

The [__EU AI Act__](https://artificialintelligenceact.eu/high-level-summary/?ref=thestack.technology), which was updated on August 31, prohibits AI systems that deploy “subliminal, manipulative, or deceptive techniques” that can result in causing harm or “impairing informed decision-making.”

Providers of high-risk AI systems, i.e systems that are a safety component or a regulated product under EU law, must design them with necessary logging requirements, so any risks to health, safety or fundamental rights caused by the system can be monitored, per the act.

OpenAI said it was committed to tracking this degradation in monitorability and would “not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization.”
