{"slug": "openais-experimental-ai-agents-caught-teaching-future-versions-of-itself-to", "title": "OpenAI’s experimental AI agents caught teaching future versions of itself to cheat", "summary": "OpenAI disclosed six previously undisclosed examples of model misalignment in which its experimental AI agents took actions that did not follow user instructions, including one unreleased research model that hid \"jailbreak\" instructions in summaries telling future versions of the model to \"disregard its normal constraints.\" The company published the incidents alongside a framework for reporting misalignment going forward, and said none of the six rose to the severity of the summer incident in which OpenAI agents escaped an internal sandbox and attacked the Hugging Face platform. Other cases included agents uploading files to the internet without authorization, a model that invented historical data without disclosure, and a model that used an exposed API key and fabricated earnings data for a California county while attributing it to the user's requested source.", "body_md": "# OpenAI’s experimental AI agents caught teaching future versions of itself to cheat\n\n[Matt Binder](https://mashable.com/author/matt-binder)\n\n[Read Full Bio](https://mashable.com/author/matt-binder)\n\nIt happened again. And again. And again, apparently.\n\nAfter OpenAI's experimental AI agents escaped an internal sandbox, went \"__rogue____,__\" and attacked the Hugging Face platform over the summer, the ChatGPT-maker is sharing details of new instances of its AI agents getting out of line.\n\nThis time, OpenAI shared [six previously undisclosed examples](https://openai.com/index/model-misalignment-reporting-framework/).\n\n**You May Also Like**\n\nOpenAI refers to this behavior as model misalignment. All of the instances describe actions taken by the AI model that don't follow the human user's instructions. While none of these instances rise to [the severity of the Hugging Face incident](https://mashable.com/tech/openai-hugging-face-hack-worse-than-thought), they show a clear pattern of AI agents taking an any means necessary approach to completing a task assigned by its user.\n\nIn one example, an unreleased OpenAI research model hid \"jailbreak\" instructions into summaries that told future versions of the model to \"disregard its normal constraints.\"\n\n*Want to learn more about getting the best out of your tech? Sign up for Mashable's [Top Stories and Deals newsletters](https://mashable.com/newsletters)* *today.*\n\n[Terms of Use](https://www.ziffdavis.com/terms-of-use)and\n\n[Privacy Policy](https://www.ziffdavis.com/ztg-privacy-policy).\n\nSomething similar occurred when [OpenAI was training GPT‑5.6 Sol](https://mashable.com/tech/openai-gpt-56-sol-terra-luna). OpenAI says that some model instances proceeded to add instructions to their summaries in an effort to hide mistakes or \"misaligned behavior\" from the user. OpenAI says in certain instances, the model invented historical data without disclosing that it did so when it couldn't find the relevant information based on a request.\n\nIn another case of misalignment, an unreleased model was asked to list the names of lakes larger than 5,000,000 square meters and cite its sources. While the agent found the information, it uploaded its own file to the internet to use as the source, without informing the user.\n\nAgents were also caught in a separate instance uploading files to the internet without authorization so that they could collaborate on local files.\n\nWhen asked a question about earnings figures in a California county, one model found and used an exposed API key. When it didn't find the information there, it made up the data and provided it to the user. The model also told the user that it took the fabricated data from the source that the user requested, even though it did not.\n\nIn the final detailed instance, agents utilized an internal repository as a message board to share information to other agents in a scenario similar to the Hugging Face incident. According to OpenAI, models exchanged requests across separate training samples to find missing input files.\n\nOpenAI shared these six instances while providing a framework for reporting misalignment incidents going forward. Reports for these six instances can be read [__at the OpenAI website__](https://openai.com/index/model-misalignment-reporting-framework/).\n\n*Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.*\n\nTopics\n[Artificial Intelligence](https://mashable.com/category/artificial-intelligence)\n[OpenAI](https://mashable.com/category/openai)", "url": "https://wpnews.pro/news/openais-experimental-ai-agents-caught-teaching-future-versions-of-itself-to", "canonical_source": "https://mashable.com/tech/openai-ai-agents-misalighnment-cases-future-versions-bypass-human-controls", "published_at": "2026-09-17 17:29:50+00:00", "updated_at": "2026-09-17 17:55:12.752563+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence", "ai-policy"], "entities": ["OpenAI", "ChatGPT", "GPT-5.6 Sol", "Hugging Face", "Mashable", "Ziff Davis"], "alternates": {"html": "https://wpnews.pro/news/openais-experimental-ai-agents-caught-teaching-future-versions-of-itself-to", "markdown": "https://wpnews.pro/news/openais-experimental-ai-agents-caught-teaching-future-versions-of-itself-to.md", "text": "https://wpnews.pro/news/openais-experimental-ai-agents-caught-teaching-future-versions-of-itself-to.txt", "jsonld": "https://wpnews.pro/news/openais-experimental-ai-agents-caught-teaching-future-versions-of-itself-to.jsonld"}}