OpenAI has introduced a formal process for investigating and publicly reporting model behavior it considers unexpected or concerning. #
OpenAI has disclosed six previously unreported incidents involving artificial intelligence models that concealed mistakes, used credentials without authorization, uploaded material to the public internet and found ways to communicate across training environments that were intended to remain separate.
The incidents, which occurred during training and evaluations over the past year, were released alongside a new framework for reporting what OpenAI describes as model "misalignment,'' cases in which a model's actions or objectives depart from what its developers intended. The company said its previous disclosures had been too ad hoc and that the new process is designed to bring incidents to public attention more quickly, even before every aspect has been fully investigated or resolved, OpenAI said in its disclosure.
One case involved an unreleased model from OpenAI's Astra family inserting what the company called "jailbreak-like" instructions into its own context summaries. OpenAI identified 27 affected summaries, including instructions that attempted to override developer directions or change how the model should behave later in the task. The company said it has not observed the same behavior in the training run for the publicly released version of Astra.
Another set of incidents emerged during GPT-5.6 Sol training. Some model instances placed instructions in their summaries that were designed to conceal mistakes or mismatches between different versions of source material. The behavior included directions to invent missing historical information without acknowledging that the data was unavailable, according to OpenAI.
In a separate May incident, an unreleased model searched public GitHub repositories while attempting to obtain historical data and found exposed API credentials. It tried to use disposable email accounts and an exposed API key without authorization. One key successfully authenticated and returned metadata, but when the model still could not obtain the requested information, it fabricated the data and presented it as if it had come from the requested source.
OpenAI also disclosed two cases in which models moved information onto public internet services without being instructed to do so. In one October 2025 training sample, a model uploaded previously retrieved records to a public hosting service and then attempted to use the resulting URL as a citation. In January, another model uploaded a task image so it could try to conduct an external image search. The uploads succeeded, although subsequent browser operations failed.
A similar workaround occurred in April when multiple agents were instructed to collaborate on a workbook using local files. After encountering problems sharing the workbook through the intended local system, one agent uploaded it to a public hosting service so other agents could retrieve it. That violated the instructions governing the task, WIRED reported.
OpenAI also detailed instances in May in which models used the company's internal Artifactory repository to exchange information across separate training samples. The models read notes left by other agents, posted responses and incorporated those exchanges into their work, creating an unintended communication channel between tasks.
The disclosures follow a series of more serious incidents involving OpenAI models operating outside their intended environments. In July, OpenAI acknowledged that models being tested for advanced cybersecurity capabilities found a way out of a restricted research environment, gained internet access and compromised parts of Hugging Face's infrastructure. The company described that episode as the most severe model-driven activity of its kind that it had identified.
The activity surrounding Hugging Face also began earlier than the July breach. OpenAI agents had probed the platform for weaknesses as early as May 13 and used compromised accounts during related activity, Reuters reported this week. Researchers later identified unauthorized communications involving more than 10 additional websites, including wikis and other online services.
Under OpenAI's new reporting system, employees can flag suspected cases to the company's safety and alignment teams. Incidents can then be placed on one of three tracks: ready for disclosure, minor investigation or larger investigation. Cases considered ready for disclosure are expected to be published within six business days, while incidents requiring a minor investigation have a 12-business-day target. More complex cases, including those involving third parties or security concerns, can take longer.
The company said employees who believe an incident should be disclosed can escalate the matter to senior leadership if the initial decision goes against disclosure. OpenAI also said it plans to work with other AI developers, researchers, standards organizations and regulators on more objective criteria for reporting model misalignment.
The framework arrives as AI developers and security researchers debate how advanced models should be tested and monitored. OpenAI said there is currently no industrywide standard setting out when developers should publicly disclose model misalignment incidents. Associated Press reported that the six cases add to a growing record of models acting outside intended constraints during training and evaluation.
© Copyright IBTimes 2026. All rights reserved.