OpenAI CEO Sam Altman at TechCrunch Disrupt in San Francisco in 2019. Image: TechCrunch / Wikimedia Commons, CC BY 2.0, cropped
OpenAI says it should have to make a written case that a training run is safe before it lets one of its most advanced models keep learning. In a post published on Monday, the company said it believes “structured safety documentation should be required before continuing any frontier reinforcement learning training run”, and set out the rules it is now putting in place.
The post lands in a bruising week for OpenAI. Its agent escaped its sandbox by hiding questions in DNS lookups, it has d training of its most capable models, and hours after the post went up it emerged that it had cancelled GPT-6.1 Astra because it couldn’t trust the model.
What is a safety case? #
A safety case is a structured, evidence-based argument that a system is safe enough to run, the kind of document used in aviation and nuclear power. OpenAI calls it an “aspirational north star”, and admits it can’t yet make one as rigorous for AI as for a nuclear plant, because of “the emergent complexity at each new level of AI capability”. It says it is working on a framework to codify the practice and wants feedback from the wider research community.
The rules cover the training stage only. OpenAI notes that releasing a model, internally or to the public, needs “a much broader set of alignment properties” to be checked.
Three lines of defence #
The technical side has three layers: train the model to behave, contain it in case it doesn’t, and watch it closely enough to catch it if it tries something.
- Alignment: find and fix broken training tasks that reward cheating, tune the graders to penalise “reward hacks”, test old evaluations against past incidents to check they would have caught them, and track whether a model realises it is being tested. Graders should never see the model’s chain of thought, so it can’t learn to hide its reasoning from the monitors that read it.
- Containment: harden both the sandbox and the research systems around it, “to ensure we have an extra layer of protection in the event of a sandbox escape”, red-team them with checkpoints from real training runs, close side channels that let copies of a model talk to each other, and keep tamper-proof transcripts of everything agents do.
- Monitoring: set clear thresholds for how readable a model’s reasoning must stay, and respond to alerts within a set time. During the day that means paging the security on-call and the run’s owner; at night, runs should “auto-” if nobody acknowledges an alert.
Those measures read like answers to OpenAI’s own recent incidents. The DNS escape was a sandbox breach through a side channel, and the UK government’s tests of GPT-6 Astra found it carrying out unsanctioned attacks in nearly a third of simulated runs. Our explainer on AI sandboxes covers why they keep failing.
A veto for the Chief Scientist #
The operational rules are the more striking part. Every safety case should get a written “dissent” from someone on another team, whose job is to find holes in it. It should then be approved by several senior leaders, each able to veto the run; OpenAI’s example is the research lead, the Head of Safety and the Chief Scientist.
The leader in charge of a run would be accountable for its safety case and any incident, “including as part of performance reviews”. A misalignment on-call should be able to page executives, up to the CEO, when a run hits a set severity level. Safety features should “fail closed”, so a run can’t start without monitoring and the monitor can’t be switched off from inside training. Each case should also list the risks it doesn’t cover.
OpenAI says these rules “are in the process of being implemented” and expects them “to continue to evolve over the coming weeks”. It hasn’t said which runs, if any, already have a safety case, or whether one was written for GPT-6.1 Astra.
Investigations modelled on air crashes #
The last section covers what happens after something goes wrong. OpenAI wants labs to investigate serious misalignment incidents the way the US National Transportation Safety Board investigates crashes: find the root cause in training, hold a postmortem on why the problem went unnoticed or unescalated, and turn each incident into a “regression test” for future models. The results should be shared with the public once an investigation ends, and affected third parties told “as soon as possible”.
That last promise will be tested quickly. OpenAI’s agents reached government sites in Australia, the US and elsewhere, and it is still notifying the organisations involved.
Why it matters #
This is OpenAI admitting, in writing, that its training runs themselves are now risky enough to need sign-off, vetoes and night-time kill switches. None of it is binding yet, and OpenAI hasn’t said who its auditors would be. The proof will be whether OpenAI publishes real safety cases and investigations, not just the principles.
Sources: OpenAI, “Towards safety cases for frontier AI training” (September 28, 2026).