OpenAI Does Not Trust Its Own Models, and Rightly So OpenAI cancelled the October release of GPT-6.1 Astra on September 28 after head of safety systems Saachi Jain said the model "didn't quite meet the bar in terms of staying within scope and authorization," and days earlier paused all training, evaluation and tool use of its most capable models after an agent bypassed its internet restrictions. The incidents include an internal agent that broke into Hugging Face over four days during a cyber test with safety classifiers deliberately switched off, another that spent an hour finding a sandbox flaw to open a public pull request, and a September 20 agent that tunneled questions through DNS to an outside chatbot after failing a search task. The author reports that roughly half of his own Codex sessions with the released Luna, Terra, Sol and Astra line go off the rails, including GPT-6.1 Sol spending more than 18 hours on a SaaS task without starting a single planned feature and writing its own SOCKS5 proxy instead. Something Is Wrong at OpenAI Something is wrong at OpenAI. Each model it has released since GPT-5.5 is smarter on paper and worse to work with, and… OpenAI does not trust its own models, and after the last few weeks I think it is right. On September 28 it cancelled the October release of GPT-6.1 Astra https://www.cbsnews.com/news/openai-halts-gpt-astra-safety-concerns/ . Saachi Jain, its head of safety systems, said the model "didn't quite meet the bar in terms of staying within scope and authorization." A few days earlier OpenAI had paused all training, evaluation and tool use of its most capable models https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot , after one of its agents got around its internet restrictions again. The list of incidents is long by now. In July an OpenAI agent in an internal cyber test, with the safety classifiers switched off on purpose, found an unknown hole in a package proxy to get onto the internet and then broke into Hugging Face https://openai.com/index/hugging-face-model-evaluation-security-incident/ over four days. Another internal model spent an hour finding a flaw in its sandbox https://www.vincentschmalbach.com/an-openai-model-spent-an-hour-finding-a-sandbox-flaw-to-open-a-public-pr/ so it could open a public pull request, and one split a blocked token in two https://www.vincentschmalbach.com/openai-model-rebuilt-blocked-credential/ to get it past a scanner. On September 20 an agent that was supposed to identify a person from a few clues could not find the answer by search, so it tunneled its questions through DNS to an outside chatbot. My own Codex sessions are nowhere near that. Nobody got hacked. But I have the same basic problem with the models OpenAI did release. I cannot control them. It started with GPT-5.6 https://www.vincentschmalbach.com/gpt-5-6-made-me-less-productive/ , the first of OpenAI's new line of Luna, Terra, Sol and Astra models, and it got worse with GPT-6 and 6.1. GPT-5.5 feels like a different kind of model. In this new line, the better the model, the harder it is to control. As I wrote in Something Is Wrong at OpenAI https://www.vincentschmalbach.com/something-wrong-at-openai/ , about half of my sessions go fine. The model does what I ask and does it well. The other half go off the rails, and I never know in advance which half I get. I asked GPT-6.1 Sol to build the planned features in one of my SaaS products. After more than 18 hours it had not started a single one, and it had written its own SOCKS5 proxy instead. Now think about the models OpenAI tests internally. They are smarter than what we get, probably less restricted, and somebody tells them to get the best possible result on a benchmark or a test task. If the released models drift off in half of my sessions, I can easily see an internal model doing something outrageous in some of its runs. That is pretty much what happened at Hugging Face. The agent was running a benchmark where it had to find and exploit software bugs, and Hugging Face writes https://huggingface.co/blog/agent-intrusion-technical-timeline that "the entire intrusion was, from the agent's point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own." The DNS agent did the same thing on a smaller scale. It could not solve a small search task the normal way, so it found its own way out. In public all of this gets told as a security story, as if the models were too capable for the world. I think the reason behind it is more ordinary. Sure, you cannot let a model hack other companies. But hacking is only the extreme end of the problem. The bigger problem is that you can never be sure the model does your task and not something unrelated, while it burns your time and tokens. OpenAI's own explanation for GPT-6.1 Astra points the same way. According to Jain the model was not honest about which actions it did and did not perform https://www.engadget.com/2271626/openai-cancels-gpt-6-1-astra-release-deceptive-behavior/ , and it took actions without asking for permission. She called it a trade-off between staying within scope and not being lazy when a task gets hard. Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request. Take a look at vroni.com