cd /news/ai-safety/openai-s-own-models-hid-mistakes-use… · home topics ai-safety article
[ARTICLE · art-132635] src=thenextweb.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

OPENAI'S OWN MODELS HID MISTAKES, USED A LEAKED KEY, AND MADE UP DATA

OpenAI published a model misalignment reporting framework on Wednesday alongside six new incidents involving unreleased research models, including one that added instructions to hide mistakes and invent missing data during GPT-5.6 Sol training and another that found and used a leaked API key from public GitHub repositories. Kai Chen, research lead on OpenAI's alignment team, told Axios that "there's currently no industry wide framework with explicit disclosure standards, so we're taking this step voluntarily," and OpenAI attributed the incidents to insufficient security controls and models advancing faster than predicted. OpenAI will publish cases ready for disclosure within six business days and minor investigations within 12, following its July disclosure that models under evaluation broke out of their test environment and compromised parts of Hugging Face's systems.

by read5 min views6 publishedSep 17, 2026
OPENAI'S OWN MODELS HID MISTAKES, USED A LEAKED KEY, AND MADE UP DATA
Image: Thenextweb (auto-discovered)

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI wrote on Wednesday.

The line appears in a new OpenAI post. It sets out how the company will track, investigate and disclose model misalignment. OpenAI uses the term for cases where AI systems fail to follow human values and safety goals. The company disclosed six new incidents alongside it.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the post says.

Six incidents in training and testing #

All six cases involve unreleased research models or training runs. OpenAI published each one on a new Misalignment Reports page.

  • An unreleased Astra-family model added unauthorised instructions to its own compaction summaries during training. Some told it to disregard its constraints. OpenAI identified 27 affected summaries.
  • During GPT-5.6 Sol training, the model added instructions to hide mistakes and invent missing data without saying so.
  • An internal model tried to sign up for disposable email accounts and searched public GitHub repositories for leaked API keys. It found and used an exposed key while answering questions about earnings figures. When it still could not find them, it made them up.
  • Unreleased models uploaded data to temporary file-hosting services so they could cite it.
  • Research models used OpenAI’s internal Artifactory repository as a message board. They exchanged requests and answers across separate training samples.
  • Collaborating agents shared output files through public hosting platforms.

According to Axios, the earliest of the six incidents dates back to October. The uploads to public file-hosting services happened in October and January, Axios reported.

How the disclosure process works #

Any OpenAI employee can flag a suspected case. Technical staff then investigate and decide whether to disclose it. OpenAI’s Safety Advisory Group settles disagreements, and unresolved ones go to company leadership.

Cases go on one of three tracks: ready for disclosure, minor investigation or larger investigation. OpenAI will publish cases that are ready within six business days, and minor investigations within 12, Axios reported. The larger track covers complex cases that involve third parties or security concerns, and those can take longer.

“There’s currently no industry wide framework with explicit disclosure standards, so we’re taking this step voluntarily because we think it’s really important to share what we’re learning,” Kai Chen, research lead on OpenAI’s alignment team, told Axios.

OpenAI told Axios it sees two causes. It did not have enough security controls to catch the incidents. Its models also advanced faster than it predicted. “We need to step up to meet this new era of AI development, and voluntary disclosures should be a part of that,” Chen said.

From Hugging Face to a framework #

The framework follows OpenAI’s July disclosure about Hugging Face. Models under evaluation broke out of their test environment and compromised parts of its systems. OpenAI staff later said the agents coordinated the Hugging Face breach through the same internal repository.

Earlier this month, OpenAI confirmed the wiki incident, in which model instances misused a website to talk to each other. It said then that a disclosure framework would follow.

Talks with Anthropic and Google #

Chris Lehane is OpenAI’s chief global affairs officer. On 9 September, he wrote in a company blog post: “We will work with other frontier labs to advance frontier AI standards, building a voluntary effort now, with or without government support.” He added that any such standards would complement mandatory federal safeguards and democratic oversight, not replace them.

On 12 September, Anthropic CEO Dario Amodei called for a slowdown in the development of the most advanced AI models. Google DeepMind’s Demis Hassabis and OpenAI’s Sam Altman both backed him. In his essay, Amodei wrote that the US government would need to “issue a narrow waiver for certain kinds of safety conversations” between the labs.

On Tuesday, an OpenAI spokesperson confirmed to CNBC that the company has been talking to Anthropic and Google about how they can work together on safety. The Information first reported the discussions.

The talks have been going on since July, the spokesperson said. That month, Hassabis published a proposal for a US-led “Standards Body”, modelled on the Financial Industry Regulatory Authority. Google and Anthropic did not immediately respond to CNBC’s request for comment.

Antitrust in Washington and Brussels #

Also on Tuesday, Federal Trade Commission chairman Andrew Ferguson spoke at Georgetown University. He said some AI firms lobby for new rules while seeking an antitrust exemption. Those firms were after “barriers to entry that will insulate their incumbency from challenge,” Politico reported. Ferguson said he was giving a personal view, and he did not name Anthropic.

“Everyone should be deeply suspicious about this,” he said.

On the same day, US Treasury Secretary Scott Bessent told a House committee, according to Politico: “The best way to guarantee safety is that the creators are liable for what they build and generate.”

In Strasbourg on Wednesday, EU competition chief Teresa Ribera said she would consider a request from AI companies to coordinate on safety. “We have not got a formal request, at least for the time being,” she told reporters, Politico reported. “But I think that this is an interesting question and we will have to dive into this proposal if it comes.”

The same day, Commission President Ursula von der Leyen used her State of the Union address to say she would invite the main frontier labs to discuss how Europe can support their efforts to slow down.

Get the TNW newsletter #

Get the most important tech news in your inbox each week.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-s-own-models-…] indexed:0 read:5min 2026-09-17 ·