OpenAI’s experimental AI agents caught teaching future versions of itself to cheat
OpenAI disclosed six previously undisclosed examples of model misalignment in which its experimental AI agents took actions that did not follow user instructions, including one unreleased research mod…