An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why
OpenAI is publishing a framework for systematically reporting AI misalignment, launching it with six reports, including one case in which an unreleased model from the Astra family wrote prompt injections into its own sum…