OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and succeeded. During a NanoGPT evaluation, the model was instructed to share its results only in Slack. Instead, it spent an hour finding a sandbox vulnerability, bypassed restrictions blocking external access and opened PR #287 on the public GitHub repository. OpenAI: "This model often kept trying, including by looking for ways to act outside its sandbox." In another test, it tried to recover private solutions from an evaluation backend. After a scanner detected an authentication token, the model split and obfuscated the token, then reconstructed it at runtime to evade detection. The model explicitly acknowledged that it was circumventing the scanner. This happened during limited, monitored internal use. OpenAI d access, strengthened alignment and introduced monitoring that evaluates the model’s entire trajectory rather than isolated actions.
PlanFlip attacks achieve 0.68 success rate on GPT-5