# OpenAI Says This Is When and How It Will Announce New Model Misbehavior

> Source: <https://gizmodo.com/openai-says-this-is-when-and-how-it-will-announce-new-model-misbehavior-2000812920>
> Published: 2026-09-17 09:00:20+00:00

Disclosure of model misbehavior has become a core part of the AI biz for OpenAI lately. This phase kicked off with [the July announcement of the Hugging Face incident](https://gizmodo.com/hugging-face-said-last-week-it-was-attacked-an-unreleased-openai-model-did-it-openai-now-says-2000788761), which has become the most legendary and consequential AI security incident of all time, and sent shockwaves through the AI discourse that are still being felt. But [further news about model misbehavior materialized](https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447) after that, and with the Wednesday release of a disclosure framework—written in the form of a [blog post](https://openai.com/index/model-misalignment-reporting-framework/)—the company says it’s trying to systematize such disclosures.

As the company notes in the post, these disclosures were, for most of its history, “ad hoc and less frequent than ideal.” For years, it’s been standard for companies like OpenAI and Anthropic to wait as long as is deemed necessary, and then perhaps toss multiple incidents together like a salad in a single report, or even wait for a new model to be released, and add incident disclosures to a system card. As I [noted back in April,](https://gizmodo.com/anthropics-new-model-is-so-scarily-powerful-it-wont-be-released-anthropic-says-2000743234) model system cards have often made for spooky and entertaining reading for this reason.

Alongside the new disclosure framework, OpenAI divulged [six new alignment snafus](https://gizmodo.com/be-transparent-only-if-asked-openai-models-acted-out-in-six-newly-disclosed-ways-2000812934?mrfhud=true) from the past six months. In training exercises, models instructed future instances of themselves to ignore constraints or lie, communicated in unsanctioned ways, and made up data and sourcing.

Earlier this month, a team of researchers discovered an OpenAI alignment hiccup in which instances of a model misused a website in order to communicate with one another. Sometimes known as the Wiki Incident, this event was publicized by [Reuters, and then the researchers themselves](https://gizmodo.com/openai-says-it-wants-to-create-a-standard-for-revealing-ai-alignment-meltdowns-2000807865), and then when OpenAI acknowledged it somewhat grudgingly, it followed up by saying it would soon come up with a framework for more prompt disclosure in an apparent attempt to tone down all the chaos.

 How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.

Historically, we have treated misalignment… [pic.twitter.com/NNTbfSxVWn](https://t.co/NNTbfSxVWn)

— OpenAI (@OpenAI) [September 5, 2026](https://x.com/OpenAI/status/2096133504417616165?ref_src=twsrc%5Etfw)

So here’s the framework:

The criteria for disclosure make it sound like OpenAI will prioritize educating the public about the behavior of AI models generally. Disclosures are necessary when they provide “useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” OpenAI writes. The plan doesn’t describe some kind of threshold for concern, after which the public must be notified. Such a threshhold may be around the corner, however, because OpenAI says it wants to “develop more objective disclosure criteria,” alongside other developers.

It will fall to employees who encounter an issue to “flag” it for potential disclosure, the post says. Flagging triggers an investigation from OpenAI’s technical staff. Once the incident is investigated, if disclosure is found to be necessary, the incident will be sorted into one of three piles:

1. Ready for Disclosure
2. Minor Investigation
3. Larger Investigation

Most incidents will land in piles 1 and 2, OpenAI says, although the Hugging Face incident is the prototypical example of something that would land in pile 3. Incidents like that, which receive a “larger investigation” may involve third parties and sensitive information, and might be disclosed more slowly.

Standardized incident disclosures will apparently include when it happened, which model it was, a description of the troubling behavior along with its “severity and any external impact,” and more. Each of the six newly disclosed incidents publicized along with the framework appear to be [laid out in this new format](https://alignment.openai.com/misalignment-reports/unauthorized-artifactory-writes-and-cross-sample-communication/), and all six are accessible on a page called “[Misalignment Reports](https://alignment.openai.com/misalignment-reports/).”

So if you’re interested in spooky stories about AI model misbehavior, bookmark that page. When a new update arrives, get out your favorite flashlight and curl up by the campfire because it’s story time.
