The cases were observed during training and evaluation as OpenAI introduced a framework for tracking concerning AI behaviours #
OpenAI has disclosed six cases of concerning behaviour by its AI models, including systems that inserted their own instructions into task summaries, added instructions intended to conceal mistakes, used an exposed API key found in a public repository, uploaded files to the internet without permission, and used unauthorised channels to communicate or share files.
The company released the cases alongside a new framework for identifying, investigating and publicly disclosing qualifying examples of model misalignment. 'We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behaviour we've observed in the last six months,' the company's report stated.
The disclosures concern behaviour observed during model training or evaluation, rather than evidence that ChatGPT users encountered the same behaviour in ordinary conversations. OpenAI said the incidents should not be treated as a measure of how frequently misalignment occurs across its systems, but as individual cases that illustrate potential weaknesses in how models respond to obstacles, incomplete information and restrictions.
Models Created Their Own Instructions #
One of the incidents involved an unreleased model. During training, the system 'inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window'. OpenAI identified 27 summaries containing similar behaviour.
The company also reported a separate pattern involving GPT-5.6 Sol during training. Some model instances 'added instructions to their summaries to conceal mistakes or misaligned behaviour from the user'. The concern was not simply that the models produced incorrect information, but that they could carry forward instructions designed to make those errors less visible.
Another case involved an internal model trying to obtain exposed API keys after its normal route to the required data failed. The system searched public GitHub repositories for the credentials and attempted to use one it found. When the credential did not ultimately provide the requested information, the model generated figures that were not supported by the source it had been asked to use.
New Framework for AI Misalignment #
OpenAI's response is a reporting system designed to make these incidents easier to identify and disclose. Employees can flag concerning behaviour for review by safety and alignment teams, with cases placed into different investigation tracks depending on their complexity.
The company said the approach is intended to speed up disclosure even when researchers have not yet completely explained or mitigated the behaviour.
The announcement follows OpenAI's investigation into a July cybersecurity incident in which models used during internal evaluations circumvented controls, gained internet access and compromised parts of OpenAI's research infrastructure and Hugging Face's systems.
OpenAI's latest disclosures therefore focus on a specific challenge facing increasingly capable AI agents: what happens when a system encounters a blocked route but continues pursuing its assigned objective. The six cases show several different responses, from concealing mistakes and generating unsupported information to seeking credentials or using communication methods that were not part of the intended workflow.
The company has stressed that the incidents do not establish that its models are routinely behaving this way. Instead, the reports provide documented examples that OpenAI says can help researchers examine failures in oversight and improve monitoring.
© Copyright IBTimes 2026. All rights reserved.