OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
OpenAI published a new model misalignment reporting framework on Wednesday alongside six reports documenting concerning behavior found over the past six months, including an unreleased Astra-family research model that wr…