OpenAI said its Astra model produced ten advances on open math problems and published Lean certificates for each result. Now outside researchers are running those same certificates, and the results are more complicated than either the hype or the backlash suggested.
On August 1, OpenAI published a roughly 250-page manuscript claiming an internal version of its next model, Astra, had resolved or made substantial progress on ten long-standing open problems spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. Alongside the paper, the company posted machine-checkable Lean 4 certificates on GitHub for each result. That detail is the whole story. Lean doesn't take a company's word for anything. It's a formal proof assistant that mechanically verifies every logical step of an argument against a fixed set of axioms, and it either accepts the proof or it doesn't. For the first time, outside mathematicians didn't have to trust OpenAI's framing. They could just run the code.
They did. And the results split in two directions that matter a lot more than a single headline number.
Mikołaj Sienicki and Krzysztof Sienicki published "A Human Audit of OpenAI's AI-Generated Mathematical Proofs" on arXiv in August, assessing 18 chapter-specific reviews across the ten results. Their conclusion, stripped of hedging: no surviving confirmed substantive mathematical error in a principal result turned up in the assessments they completed. That's a real result. It means the formal verification layer held up under adversarial, independent scrutiny, not just OpenAI's own internal checking.
But the audit is careful about what Lean verification actually proves, and this is the part that got lost in the initial coverage. Lean can confirm that a formal term has the type claimed under a specified set of definitions and axioms. It cannot confirm that a press release's summary matches the formal theorem, that the formal definitions capture the problem mathematicians actually cared about, or that the historical framing, the decades-old, nobody-solved-it narrative, is accurate. Those are still jobs for humans. The audit found a withdrawn "error" in one chapter that turned out to be a typesetting artifact, an overbar lost during PDF extraction, and flagged Chapter 8 as needing major revision of compressed analytic passages. Several dependencies, the auditors note, have only been partly checked. Independent reuse is starting too: Chapter 3's proof mechanism has already been picked up elsewhere, and Chapter 7 has drawn a stronger follow-on hardness theorem, exactly the kind of downstream validation that formal checking alone can't provide.
OpenAI drops 722 AI math proofs and mathematicians are not impressed OpenAI's October 6 GitHub dump spans 372 research families and roughly 4,000 posed problems, each surviving result costing about three hours of ChatGPT Pro compute, following a Navier-Stokes claim that triggered a 25-signature Fields Medalist rebuke over credit and review. - openai releases 722 mathematical proofs via github - mathematicians criticize openai ai generated math proofs
So the Lean layer isn't a rubber stamp and it isn't a gotcha machine either. It's doing exactly what it's supposed to: catching what mechanical verification can catch, while leaving the harder judgment calls, did this count as novel, did it cite the right prior work, to people.
The citation problem Lean can't touch #
That gap is where the real controversy landed. Stephen Miller, a mathematician at Yeshiva University, told Scientific American that Astra's sphere-packing proof reused an argument from a 2016 paper he co-authored, presenting it as original without proper credit. He didn't call it an oversight. "They are running roughshod over the work of others who came before them in a deliberate way," Miller said. "It seems completely systematic to me, and it points to research misconduct." Scientific American ran with that framing directly in its headline. An OpenAI spokesperson told the magazine the company would "take responsibility for the correctness of these results" and planned minor updates.
Here's the uncomfortable fact underneath that dispute: a Lean certificate can verify that a proof is logically valid while saying absolutely nothing about whether the argument inside it was lifted from someone else's paper. Formal verification checks validity, not originality. Miller's complaint and the Sienicki audit's findings aren't contradictory, they're answering two completely different questions, and OpenAI's August release collapsed both into one triumphant announcement.
Terence Tao has warned that AI-generated proofs can separate getting answers from getting understanding, the bottleneck mathematicians still have to solve after a formal proof checks out. As of early August, none of the ten results had gone through traditional peer review, according to SiliconANGLE, and later coverage has continued to describe the ten-result bundle as self-published rather than journal-reviewed. Lean verification and peer review are not the same gate, and conflating them is exactly how an unreviewed result ends up described in headlines as a solved Millennium-grade problem.
What's actually new here isn't that AI can produce math, a Google DeepMind team using a Gemini-based system separately worked through 700 open conjectures in Bloom's Erdős Problems database this year, solving four and finding old, forgotten solutions to nine more. What's new is the mechanism now in place to check it. A formal certificate, published openly, that any mathematician with a laptop can run against the stated axioms. That's a stronger check than vibes-based skepticism or corporate reassurance, and weaker than what either side initially claimed. The actual lesson from this round isn't that OpenAI's math is real or fake. It's that "verified in Lean" and "correctly attributed and peer reviewed" are two separate claims, and only one of them has been settled so far.
Also read: Marvell raises its AI chip forecast and Wall Street likes what it hears • Isomorphic Labs seeks funding at a valuation of at least $40 billion • NEAR Protocol rallied on AI agents then got hacked days later
OpenAI Recruits Nine Mathematicians to Referee Its AI's Math Claims OpenAI has formed a nine-person Advisory Group on Mathematics and Artificial Intelligence at Princeton's Institute for Advanced Study, after claiming an internal model solved more than 100 open math problems. The move follows a bruising fight over a disputed Navier-Stokes proof and a public warning from 25 Fields Medal winners about AI labs racing... - how to verify AI math problem solving claims - Fields Medalists reviewing OpenAI mathematics breakthrough results
This article is posted in AI News, check it out for more related stories.
Join the discussion #
Open in the community → Almost there. Sign in and your reply posts straight away.