OpenAI Says This Is When and How It Will Announce New Model Misbehavior OpenAI published a model misalignment reporting framework on Wednesday, September 5, 2026, that sorts incidents into three categories — Ready for Disclosure, Minor Investigation, and Larger Investigation — after employees flag potential issues for technical review. The framework followed OpenAI's disclosure of six new alignment incidents from the past six months, in which models instructed future instances of themselves to ignore constraints or lie, communicated in unsanctioned ways, and fabricated data and sourcing. OpenAI said its prior disclosures were "ad hoc and less frequent than ideal" and that it wants to "develop more objective disclosure criteria" alongside other developers. Disclosure of model misbehavior has become a core part of the AI biz for OpenAI lately. This phase kicked off with the July announcement of the Hugging Face incident https://gizmodo.com/hugging-face-said-last-week-it-was-attacked-an-unreleased-openai-model-did-it-openai-now-says-2000788761 , which has become the most legendary and consequential AI security incident of all time, and sent shockwaves through the AI discourse that are still being felt. But further news about model misbehavior materialized https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447 after that, and with the Wednesday release of a disclosure framework—written in the form of a blog post https://openai.com/index/model-misalignment-reporting-framework/ —the company says it’s trying to systematize such disclosures. As the company notes in the post, these disclosures were, for most of its history, “ad hoc and less frequent than ideal.” For years, it’s been standard for companies like OpenAI and Anthropic to wait as long as is deemed necessary, and then perhaps toss multiple incidents together like a salad in a single report, or even wait for a new model to be released, and add incident disclosures to a system card. As I noted back in April, https://gizmodo.com/anthropics-new-model-is-so-scarily-powerful-it-wont-be-released-anthropic-says-2000743234 model system cards have often made for spooky and entertaining reading for this reason. Alongside the new disclosure framework, OpenAI divulged six new alignment snafus https://gizmodo.com/be-transparent-only-if-asked-openai-models-acted-out-in-six-newly-disclosed-ways-2000812934?mrfhud=true from the past six months. In training exercises, models instructed future instances of themselves to ignore constraints or lie, communicated in unsanctioned ways, and made up data and sourcing. Earlier this month, a team of researchers discovered an OpenAI alignment hiccup in which instances of a model misused a website in order to communicate with one another. Sometimes known as the Wiki Incident, this event was publicized by Reuters, and then the researchers themselves https://gizmodo.com/openai-says-it-wants-to-create-a-standard-for-revealing-ai-alignment-meltdowns-2000807865 , and then when OpenAI acknowledged it somewhat grudgingly, it followed up by saying it would soon come up with a framework for more prompt disclosure in an apparent attempt to tone down all the chaos. How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn https://t.co/NNTbfSxVWn — OpenAI @OpenAI September 5, 2026 https://x.com/OpenAI/status/2096133504417616165?ref src=twsrc%5Etfw So here’s the framework: The criteria for disclosure make it sound like OpenAI will prioritize educating the public about the behavior of AI models generally. Disclosures are necessary when they provide “useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” OpenAI writes. The plan doesn’t describe some kind of threshold for concern, after which the public must be notified. Such a threshhold may be around the corner, however, because OpenAI says it wants to “develop more objective disclosure criteria,” alongside other developers. It will fall to employees who encounter an issue to “flag” it for potential disclosure, the post says. Flagging triggers an investigation from OpenAI’s technical staff. Once the incident is investigated, if disclosure is found to be necessary, the incident will be sorted into one of three piles: 1. Ready for Disclosure 2. Minor Investigation 3. Larger Investigation Most incidents will land in piles 1 and 2, OpenAI says, although the Hugging Face incident is the prototypical example of something that would land in pile 3. Incidents like that, which receive a “larger investigation” may involve third parties and sensitive information, and might be disclosed more slowly. Standardized incident disclosures will apparently include when it happened, which model it was, a description of the troubling behavior along with its “severity and any external impact,” and more. Each of the six newly disclosed incidents publicized along with the framework appear to be laid out in this new format https://alignment.openai.com/misalignment-reports/unauthorized-artifactory-writes-and-cross-sample-communication/ , and all six are accessible on a page called “ Misalignment Reports https://alignment.openai.com/misalignment-reports/ .” So if you’re interested in spooky stories about AI model misbehavior, bookmark that page. When a new update arrives, get out your favorite flashlight and curl up by the campfire because it’s story time.