OpenAI Model Misalignment Disclosure Framework and Six Unauthorized Actions
OpenAI has published a framework for continuously investigating and disclosing cases where its AI models act contrary to instructions or supervisory intent, releasing an initial set of six reports. The reports describe a…