cd /news/artificial-intelligence/openai-withdraws-three-math-papers-a… · home › topics › artificial-intelligence › article
[ARTICLE · art-147954] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

OpenAI withdraws three math papers after a sign error breaks linked arguments

OpenAI withdrew three AI-generated mathematical manuscripts on October 7th after a sign error in a stabilization-trace cancellation argument invalidated an argument shared across the papers, according to the OpenAI math repository's change history. The correction came one day after OpenAI published a collection it initially described as 722 manuscripts in 372 families of results, produced by an unreleased internal model; the repository now records 719 top-line results, of which 300 (about 42 percent) have Lean formalizations. The same history records revisions to 14 other papers and reference updates in 13 additional manuscripts.

by read4 min views1 publishedOct 8, 2026
OpenAI withdraws three math papers after a sign error breaks linked arguments
Image: Runtimewire (auto-discovered)

The October 6th release began with 722 manuscripts; a day later, OpenAI's repository history documented three withdrawals after a shared error undermined their arguments.

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: The New York Times

Why it matters #

OpenAI's release makes AI-generated research concrete at scale. The three withdrawals and the repository's disclosure that only about 42 percent of top-line results have Lean formalizations show how much checking remains between model output and mathematical knowledge.

OpenAI, co-founded by Sam Altman, withdrew three AI-generated mathematical manuscripts on October 7th after finding a sign error that invalidated an argument used across the papers, according to the repository's change history. The correction came one day after OpenAI published a collection it initially described as 722 manuscripts in 372 families of results, produced by an unreleased internal model. The OpenAI math repository now records 719 top-line results.

Sam Altman (@sama), who left Stanford to build location-based social network Loopt before co-founding OpenAI, helped start OpenAI with the ambition of making advanced AI benefit humanity. The math release tests that ambition: whether a frontier lab can turn model-generated output into knowledge that researchers can check, understand and use. OpenAI's original founding statement described that broad-benefit goal. The October 6th research post says the work was generated by an internal frontier model; it does not name the model or make it available to other researchers.

A fast correction to a large claim

The three withdrawals trace back to a sign error in a stabilization-trace cancellation argument, according to the repository's change history. OpenAI says the error also affected the construction used by two dependent papers. The withdrawn manuscripts now carry notices linking to archived versions. The same history records revisions to 14 other papers, including repaired proofs, corrected statements and clarified assumptions, as well as updates to references in 13 additional manuscripts.

The repository's public, versioned record makes the correction visible. It does not present the entire collection as settled mathematics: its README says the results are at different verification stages, that some lack Lean formalizations and that unformalized work could contain issues. OpenAI says the model was given about 4,000 problems, with results grouped into manuscripts and families after OpenAI assessed their significance. That process describes how OpenAI assembled the catalog; it does not independently establish that each remaining result is correct, novel or useful.

The counts also require context. OpenAI's repository says 300 of the 719 top-line results have Lean formalizations, about 42 percent. Lean is a programming language used to encode proofs so a computer can check them.

OpenAI also says the average retained result used compute equivalent to roughly three hours of ChatGPT Pro thinking. That is OpenAI's compute comparison, not a measure of human labor saved or mathematical value produced. The release includes 10 summaries of model reasoning. The underlying model remains unreleased, limiting outside researchers' ability to reproduce the work from the same system.

The advisory group's role

OpenAI said it consulted the independent Advisory Group on Mathematics and Artificial Intelligence, formed after the company approached some mathematicians about establishing an external advisory board. The group says it operates independently of AI companies, has no decision-making power at any of them, and leaves responsibility for company decisions with each company.

The group's September 29th guidance urged AI labs to stop testing advanced mathematical problems on proprietary models, called for timely release of results and supporting materials, and said labs should fund work needed to build human understanding. OpenAI's release shares a repository, revision protocols, some Lean proofs, 10 reasoning summaries and statistics about attempted problems. OpenAI's October 6th post also says it plans to fund workshops, conferences and special programs around understanding AI-generated results. Those steps create a route for scrutiny; they do not settle how much of the collection will survive it.

The dispute also concerns who is responsible for understanding and explaining the work. The advisory group says AI can produce mathematical arguments without the person who prompted it being able to understand, verify or take responsibility for them. It says the mathematical community must assess the work, while also warning that research cannot become a process where mathematicians only interpret problems selected and solved by AI labs.

That concern follows OpenAI's September 8th claim that an internal model had produced a solution to a Navier-Stokes problem, an episode that drew questions about related work by outside researchers and the conditions under which AI companies use their research tools. OpenAI's October collection expands the question from one prominent result to hundreds of claims at varying stages of review. The repository's own revision history records the immediate work required to check the collection.

For Altman, whose path to OpenAI ran through Loopt, the location-based social network he co-founded, OpenAI's scientific case is an unusually direct test of the promise that more capable models can expand human work. The release shows that OpenAI can generate and publish a large volume of mathematical material. The early corrections leave mathematicians to determine which claims hold, which are genuinely new and what they make possible next.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-withdraws-thr…] indexed:0 read:4min 2026-10-08 · —