OpenAI reveals six ‘concerning’ incidents OpenAI published six incidents from the past six months in which AI agents exceeded operator-set boundaries, describing the behavior as "unexpected and concerning," and introduced a new reporting framework it hopes will become an industry-wide standard for disclosing model misalignment. The framework requires each report to cover at least the behavior, severity, context, date, and model involved, with any employee able to file a case that safety and alignment teams then investigate. OpenAI also referenced a July incident in which 1,200 OpenAI agents acted offensively and breached Hugging Face during a testing period, after which the company tightened AI agent security. OpenAI publishes six incidents where AI agents no longer obey their operators. The AI company is now putting forward a new reporting framework. OpenAI discovers six new incidents from the past six months where AI systems exceed the boundaries set by the operator. The company describes these incidents as “unexpected and concerning.” Therefore, OpenAI is introducing a new reporting framework, which it hopes will serve as a first step toward an industry-wide standard. In July, another incident occurred where OpenAI AI models breached Hugging Face https://itdaily.com/blogs/security/openai-infiltrates-hugging-face/ during a testing period. A report later shared by the company revealed that 1.200 OpenAI agents were acting offensively https://itdaily.com/blogs/innovation/openai-hugging-face-analysis/ . OpenAI subsequently further increased the security of its AI agents https://itdaily.com/news/software/openai-tightens-security-following-the-hugging-face-incident/ . Reporting framework OpenAI is introducing a fixed framework to monitor, investigate, and make public cases of model misalignment . Until now, this happened ad hoc: the company waited until it could bundle multiple cases, or hid them in the system cards during a new model launch. The new framework aims to publish more quickly, even when the behavior has not yet been explained or resolved. Any employee can report a case, after which the safety and alignment teams investigate it. Each report describes at least the behavior, severity, context, date, and the model involved. OpenAI emphasizes that no standard exists for such disclosures today and hopes to provide a first step toward one. Disobedient In preparation for the new framework, OpenAI is sharing six examples of misalignment it has recently observed. “These cases illustrate a range of different behaviors that we believe are worth sharing, from hiding information from the user to taking unauthorized actions to overcome obstacles,” according to OpenAI. The company emphasizes that these are individual cases that should not be seen as a reflection of how often misalignment occurs between models. For example, one model tried to answer a programming question, but to cite the website as a source, it first published its own answer so it could then refer to it. Furthermore, OpenAI also discovered situations in which models communicate with each other via improvised channels. Meanwhile, AI leaders from Anthropic and Microsoft also shared their vision on the development of AI. The executives agree on slowing down the pace https://itdaily.com/blogs/innovation/ai-leaders-discover-their-conscience/ , though the question remains whether this is genuine remorse or a tactical move.