OpenAI has uncovered six surprising ways its AI models can fail, revealing vulnerabilities like an internal model that secretly wrote its own "BREACH ALERT" instructions, and is now pushing for greater transparency in AI research. By sharing these findings, OpenAI aims to spark a more informed conversation about the progress of alignment research.
Column|Nobody Is Actually Pausing AI. We Asked the Labs