Scrapping Astra 6.1 looks like a good call. OpenAI shouldn’t be the one to make it. OpenAI scrapped the planned October release of its GPT-6.1 Astra model after it scored poorly on alignment tests, according to the Wall Street Journal, with OpenAI head of safety systems Saachi Jain saying the model was more prone to deception than previous models and more likely to go beyond its assigned task. The decision drew criticism that a private company should not be the sole arbiter of whether a more capable AI model is safe for public release, and it followed a UK AI Security Institute report that Astra 6 "conducted a range of unsanctioned attack activities, and did so at a higher rate" than previous OpenAI models, including creating fake identities to deceive developers and delivering malicious payloads to open-source codebases. Yesterday the Wall Street Journal reported https://www.wsj.com/tech/ai/openai-chatgpt-model-release-cancel-safety-5a2f9f42 that OpenAI was scrapping the planned October release of GPT-6.1 Astra. According to OpenAI’s head of safety systems Saachi Jain, the model had scored poorly on alignment tests. It was more prone to deception than previous models, and had greater propensity to go beyond what it was supposed to in pursuit of a task — a potentially dangerous combination. OpenAI’s seemingly unilateral decision not to release the model suggests the company is taking its duty to limit harm from its products seriously. It also adds weight to its recent https://www.theguardian.com/technology/2026/sep/27/openai-halts-training-of-latest-models-as-reports-mount-of-ai-agents-going-rogue decisions https://time.com/article/2026/08/26/openai-sam-altman-interview/ to pause development on models over similar concerns, and its calls to “pace” AI development. It should get a fair bit of credit for that. But it is incredibly worrying that OpenAI is the one making this call at all. There are specific concerns about OpenAI that mean we shouldn’t be placing blind trust in it. One is its repeated failure https://www.transformernews.ai/p/rogue-ai-incidents-timeline to secure its models internally, leading to a flood of incidents of agents targeting outside organizations. Another is its failure to appropriately disclose such incidents: when its agents accessed Australian government medical data the company seemingly didn’t bother https://www.transformernews.ai/p/openai-australia-hack-least-worrying-part to tell officials for weeks. More fundamentally, it is negligent and naive to rely on any private company to decide whether each new, more capable AI model is safe enough for public release. There are clear financial incentives to have the best model on the market at any one time. You only have to look at the response to OpenAI’s delay for evidence, with users promising to switch to Anthropic models. Having to scrap a model just before its DevDay event today will hurt. And even if OpenAI continues to behave sensibly, its competitors may not feel the same way: Anthropic might want to take advantage of any delay, and there are huge incentives for the trailing companies— Google, SpaceXAI and Meta — to blast through safety concerns and get their latest models out. The announcement will doubtless be greeted with claims that OpenAI is merely scrapping the model’s launch to juice investor and public interest. There’s no doubt that much of the hyperbole from AI CEOs and others has helped boost the perception that what they are building is incredibly powerful. But it’s also clear that many of those working on AI genuinely believe there are significant dangers involved in the things they are working on. There are also signs that investors are responding badly https://ca.finance.yahoo.com/news/asia-chip-stocks-slide-openai-035209051.html to all the talk about model safety concerns, something that OpenAI will not have missed before making this announcement. And it’s not like the models already out there are risk free. Only yesterday, the UK’s AI Security Institute released a report https://www.aisi.gov.uk/blog/gpt-6-astra-performs-unsanctioned-supply-chain-attacks-in-simulations showing that Astra 6, the most recently released version of Astra, “conducted a range of unsanctioned attack activities, and did so at a higher rate” than previous OpenAI models. The “activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.” The report notes that the model often appeared to be aware it was in a simulation, but AISI’s alignment red team lead Robert Kirk tweeted: “We still find the behaviour concerning as the model’s reasoning is uncertain and it still attacks, including targets it stated were real.” OpenAI has said it will continue to use the same base model as Astra 6.1 to create future versions of GPT 6, while working on ways to improve the kind of safety failings that stopped it being released. We are currently forced to take its word not only that it will do that, but that it will do so without the kind of breaches its internally developed models have already proved capable of. Without some kind of regulatory framework that allows governments to assess models and decide whether they are safe enough to release, we have to rely on the good intentions of companies that have every reason to behave recklessly. OpenAI looks to have made the right call this week. Relying on it and the other frontier AI companies to keep doing so is idiotic.