This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment
While generalization is hard and I worry about our metrics being gamed we are putting a lot of effort into general solutions rather than adding alignment datasets Improvements for Astra came from more general techniques in development long long before the Hugging Face incident. ExploitGym Honeypot was very recently added as an eval following that incident and is out of distribution for our RL runs. There are also clear improvements across
GPT-6 is significantly better aligned than 5.6 but less monitorable. It is our first model to evade CoT-only monitors in sabotage evals and can sandbag without detection (which it feels like sometimes does). Hopefully we can reverse this trend.