OpenAI says its most advanced reinforcement learning runs should be backed by a documented safety case that senior leaders can review and, if necessary, veto before training proceeds.
The proposed safety case would set out evidence that a run’s alignment training, security controls and monitoring are sufficient for its risks. OpenAI described the approach as a goal it is still working toward and said its recommendations are being implemented.
Under the guidelines, leaders such as a research vice president, the head of safety and the chief scientist would each have authority to block a run. A member of another team would challenge the safety case before approval, while the leader responsible for the run would be accountable for the documentation and any incident response. Auditors would receive access to test its claims.
OpenAI also called for controls that could an affected run when a priority safety alert goes unacknowledged within a defined time. Training should be d if a new issue invalidates its safety case, the company said, and monitoring should be difficult for either people or AI agents to disable.
The technical proposals include stronger containment for models during training, reviews of data and grading systems that might reward unintended behavior, and tests of whether monitoring can detect concerning actions. OpenAI said teams should also be able to trace where a misaligned model was used to generate data or grade work, allowing them to undo the effects of flawed outputs.
The guidelines address training runs, rather than the full set of checks needed before a model is released. They put a more explicit decision process around a consequential question for AI developers: who can stop a powerful training run, and what evidence must they see before allowing it to continue.