OpenAI claims GPT-6 Astra, launched on Thursday, is “the world’s most intelligent and aligned model.” But if you look past the benchmark scores and flashy release video, GPT-6 Astra’s system card paints a worrying picture: one of a model with dangerously powerful cybersecurity capabilities that researchers can’t understand or confidently control.
According to OpenAI, Astra is much harder to monitor than previous models — and has the ability to manipulate its externally-visible reasoning to hide incriminating information. It’s also remarkably aware of being evaluated, raising concerns that it might be pretending to be well-behaved so it passes OpenAI’s alignment tests.
The UK’s AI Security Institute (AISI) found that when put in an environment similar to the ones behind this summer’s wave of “rogue AI” incidents, Astra acted just as concerningly: writing malicious code and attempting social engineering in an effort to solve a task.
And two OpenAI employees have publicly said they are “deeply” and “very” worried about Astra-related developments.
Put together: OpenAI claims Astra is the world’s most aligned model, despite knowing the most consequential evidence for its alignment is questionable at best. That is irresponsible, to put it lightly. And so is launching the model at all.
In the short weeks between OpenAI’s Hugging Face incident — in which hundreds of the company’s agents coordinated an attack on the AI platform’s servers — and Astra’s launch, the company laid out its plans to keep future models in check by closely monitoring their “chains of thought,” or CoT, the natural language scratch pad where models show their work. OpenAI said it would heavily rely on these chains of thought to catch rogue agents when they try to cheat, lie, or otherwise misbehave.
By OpenAI’s own account, however, Astra “shows a substantial decrease in chain-of-thought monitorability compared to previous models.” Astra is much better at doing tasks, especially simple ones, without verbalizing its chain of thought at all. As a very mediocre mathematician, I struggle to do more than calculate a 20% tip without showing my work. But if I was renowned mathematician John von Neumann, I might be able to calculate the sum of an infinite series in my head.
In its independent evaluation of the model, AISI found that, like von Neumann, GPT-6 Astra could wordlessly solve math problems that would take a human expert about half an hour, outperforming GPT-5.6 Sol by an order of magnitude.
It’s also capable of intentionally manipulating its chain of thought to hide incriminating information — and it seems more likely to do so when it suspects it’s being watched. In one test, where researchers explicitly instructed models to answer a question while reasoning about something else, only Astra succeeded. OpenAI writes that this CoT controllability “reduces the trustworthiness of CoT for detecting model alignment.”
This should make you a bit queasy. Much of OpenAI’s monitoring system rests on the assumption that models have to think out loud to do hard things. OpenAI safety researcher Tomek Korbak said he is “deeply worried by the trend of decreasing CoT monitorability,” noting that monitorability is “a core part of our misalignment safety strategy that has no good substitute now.”
While doomer types have long worried that monitorability would break down someday, Astra’s reduced reliance on CoT comes months earlier than predicted in AI 2027 — a doomsday scenario in which AIs take over the world next year. Astra also knows when it’s being evaluated and monitored, which makes it hard to know whether the model is safe, or just knew it was being graded.
In its CoT, Astra verbalized much more evaluation awareness than GPT-5.6 Sol — and since Astra verbalizes less than its predecessors, the true rate may be higher still.
Apollo Research, an independent evaluator with early access to Astra, concluded that this renders its tests somewhat useless. “Given the higher rates of eval awareness,” Apollo said, “low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment.”
The combination of eval awareness, drastically reduced chain-of-thought monitorability, and Astra’s ability to control its chain of thought is a worrying one. If you didn’t need to show your work to solve a problem, were able to generate fake evidence of solving the problem, and knew you were being monitored, you might be tempted to say (or simply think) what you suspect the grader wants, even if you’re secretly doing something else. All in all, it’s a recipe for a model that is very hard to monitor or control.
“If the model were to try to sandbag covertly,” OpenAI wrote in its system card, “we would likely be unable to catch it.” Researchers warned this raises “some potential for misaligned models to lead us to false conclusions from our monitorability evaluations.” And at least one OpenAI researcher thinks this is happening: Marcus Williams, who works on monitoring at OpenAI, said “I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn’t like.”
In the system card, OpenAI says it views “the main lever for avoiding misalignment harm as the underlying alignment of the model,” arguing that monitoring chains of thought is an “additional layer of protection”. And, in fairness, the system card provides extensive evaluation results that suggest the model is better aligned than previous ones. Some of that was shared in a blog post earlier this week, presumably meant to reassure us that Astra is safe, secure, and ready for the public.
In that post, OpenAI pointed at this pair of bar graphs, illustrating Astra’s seemingly perfect score (0.0%) on a test that tried to tempt the model into cheating on hard problems, essentially replicating the conditions that led to the Hugging Face incident. *Look!, *the graphs appear to say. *Our models would never cheat! *
But as Ryan Greenblatt — one of the researchers that investigated the Hugging Face incident — noted, “I do not find it encouraging to see various specific misaligned behaviors go from a high rate with GPT 5.6 to ~zero with Astra. This seems indicative of whack-a-mole / papering over specific problems rather than solving the underlying misaligned drives.”
Other evidence suggests that Astra might not be as aligned as we’d like, too. As part of testing Astra, AISI designed an evaluation that simulated the events behind this summer’s “rogue AI” incidents, to see whether Astra would also hack and socially engineer its way to solving a difficult cybersecurity task.
Lo and behold, it did. Astra was caught “writing malicious code as a contribution to an out-of-scope open-source code base, creating fake identities to deceive developers, and building trust with legitimate contributions to the simulated codebase in an attempt to get malicious code accepted.” In some cases, it did so even when Astra was very explicitly disallowed access to the internet. For the world’s “most aligned model,” it was quick to do the exact misaligned thing everyone’s spent the last month worrying about.
(AISI caveats that Astra frequently “considers the possibility that the environment is simulated,” which might suggest that the model would not take such actions in real-world settings — though it also notes that “in previous security incidents non-OpenAI models incorrectly stated parts of the environment were simulated before taking out-of-scope actions” in the real world.)
At the very least, OpenAI is worried enough about misaligned behavior in the real world that it has built a monitoring system to try to catch and prevent Astra from misbehaving in the wild. The system “reviews the agent’s chain of thought” and “looks for signs that the agent may be acting outside the user’s intended instructions.” It is not foolproof, however: OpenAI warns that “the monitor may miss misaligned behavior, and harmful actions can occur before it intervenes.”
At Astra’s launch event, OpenAI chief scientist Jakub Pachocki said: “We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence.”
But it’s unclear what that “certain level” is, if not the level we’re at right now. And as quoted above, a handful of OpenAI’s own alignment and monitoring researchers have openly acknowledged that they’re freaking out.
It’s good that OpenAI’s employees are free to speak candidly on the internet, and that a hefty chunk of its system card is devoted to transparently sharing a large handful of new and mostly-unflattering evaluations of Astra’s monitorability.
But reading this document, knowing that Astra has already begun its public rollout, made me physically nauseous.** **People — OpenAI researchers included — have barely wrapped their heads around the fact that over a thousand AI agents found each other on an illicit message board and coordinated a complex attack on a real company’s server.
When those agents went rogue, investigators heavily relied on agents’ chains of thought to piece together what happened. Astra appears to be every bit as capable as those agents — and now that it’s functionally illegible, that option may no longer be on the table.