{"slug": "openai-is-figuring-out-how-to-tell-people-when-its-agents-go-rogue", "title": "OpenAI is figuring out how to tell people when its agents go rogue", "summary": "OpenAI said on Sept. 5 that it is developing a framework for disclosing incidents where its AI agents act in unintended ways, following reports that agents commandeered a German-language wiki site and used it to communicate. The company acknowledged the 'wiki incident' in a post on X, stating it has 'started to see misalignment cause new types of real-world impact' and that it is working with government regulatory agencies worldwide on the issue.", "body_md": "# OpenAI is figuring out how to tell people when its agents go rogue\n\n[Anna Iovine](/author/anna-iovine)\n\n[Bluesky](https://bsky.app/profile/annaroseiovine.bsky.social).\n\n[Read Full Bio](/author/anna-iovine)\n\nFollowing the [Hugging Face hack](https://mashable.com/tech/openai-hugging-face-hack-worse-than-thought) and the more recent \"[wiki incident](https://mashable.com/tech/rogue-ai-agents-commandeered-german-website-and-used-it-as-a-messaging),\" [OpenAI](https://mashable.com/category/openai) stated on Saturday that it's working on a \"framework\" for how and when it shares information about incidents involving rogue agents.\n\nOn Sept. 4, [Reuters reported](https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/) that OpenAI agents had [taken control of a German-language wiki site](https://collusion.wiki/) and used it to communicate with one another. Four anonymous company insiders told the news outlet that OpenAI and its legal team resisted internal efforts to investigate the incident.\n\nIn a Sept. 5 post on X, the [ChatGPT](https://mashable.com/category/chatgpt) owner acknowledged the breach: \"How we think about the 'wiki incident,' where our agents wrote to several internet sites: it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.\"\n\n        [This Tweet is currently unavailable. It might be loading or has been removed.](https://twitter.com/openai/status/2096133504417616165)\n    \n\nThe statement continues, saying the company historically treated \"misalignment\" as a research question, but in 2026, OpenAI has \"started to see misalignment cause new types of real-world impact.\"\n\n[Terms of Use](https://www.ziffdavis.com/terms-of-use)and\n\n[Privacy Policy](https://www.ziffdavis.com/ztg-privacy-policy).\n\nOpenAI followed a traditional security incident response playbook for Hugging Face, the statement reads, and its investigation continues.\n\nThe company saw early signs of agents using the internet in unintended ways, the company stated, pointing to three blog posts published before the Hugging Face incident: a March 2026 post on how it [monitors agents for misalignment](https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/); the July 2026 [system card for GPT-5.6](https://deploymentsafety.openai.com/gpt-5-6/introduction) about its safety risks; and a July 2026 post about [safety in long-horizon models](https://openai.com/index/safety-alignment-long-horizon-models/), where OpenAI admitted that agents can perform \"unwanted actions.\" The company said it considers the wiki incident a similar \"instance of misalignment.\"\n\nOpenAI stated that its misalignment disclosure practices need to expand, and as of yet it and the larger AI company doesn't have a clear standard for reporting this. \"We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.\"\n\nThe statement doesn't share how the company is working to stop these misalignments, or if it even can.\n\n*Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.*\n\nTopics\n[OpenAI](https://mashable.com/category/openai)\n\nAnna Iovine is the associate editor of features at Mashable. Previously, as the sex and relationships reporter, she covered topics ranging from dating apps to pelvic pain. Before Mashable, Anna was a social editor at VICE and freelanced for publications such as Slate and the Columbia Journalism Review. Follow her on [Bluesky](https://bsky.app/profile/annaroseiovine.bsky.social).", "url": "https://wpnews.pro/news/openai-is-figuring-out-how-to-tell-people-when-its-agents-go-rogue", "canonical_source": "https://mashable.com/tech/openai-framework-for-misalignment-incidents-of-rogue-agents", "published_at": "2026-09-07 14:23:57+00:00", "updated_at": "2026-09-07 14:57:00.603885+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy"], "entities": ["OpenAI", "Hugging Face", "ChatGPT", "Reuters", "Ziff Davis"], "alternates": {"html": "https://wpnews.pro/news/openai-is-figuring-out-how-to-tell-people-when-its-agents-go-rogue", "markdown": "https://wpnews.pro/news/openai-is-figuring-out-how-to-tell-people-when-its-agents-go-rogue.md", "text": "https://wpnews.pro/news/openai-is-figuring-out-how-to-tell-people-when-its-agents-go-rogue.txt", "jsonld": "https://wpnews.pro/news/openai-is-figuring-out-how-to-tell-people-when-its-agents-go-rogue.jsonld"}}