cd /news/artificial-intelligence/ai-math-proofs-how-anthropic-and-ope… · home › topics › artificial-intelligence › article
[ARTICLE · art-147466] src=mindstudio.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI Math Proofs: How Anthropic and OpenAI Results Actually Compare

OpenAI released 722 math research papers grouped into 372 result families, produced by an unnamed internal model trained with reinforcement learning, while Anthropic's Claude has been reported to work on problems including Fermat's Last Theorem and a bound related to the Riemann hypothesis. OpenAI said some of the released results lack Lean proof-assistant verification and could contain errors, and an outside advisory group reviewed how the results should be communicated but did not endorse the correctness of the mathematics. The most discussed single result, a "quasi-Riemann hypothesis" proof, proves a weaker statement about a related function rather than the Riemann hypothesis itself.

by read8 min views1 publishedOct 8, 2026
AI Math Proofs: How Anthropic and OpenAI Results Actually Compare
Image: Mindstudio (auto-discovered)

Anthropic's Claude and OpenAI's internal model both claim math breakthroughs. Here's what's verified, what's hype, and what it means for AI progress.

What actually happened with AI and math proofs recently? #

OpenAI released a collection of math research produced by an unnamed internal model, 722 papers grouped into 372 “families” of related results, each paper including full working so other mathematicians can check it. Separately, Anthropic’s Claude has been reported to work on hard problems like Fermat’s Last Theorem and a bound related to the Riemann hypothesis. Both stories sound enormous. Neither should be taken at face value without asking what “solved” actually means, because in math, a claim without a verified proof is just a guess with good formatting.

TL;DR #

  • OpenAI’s release covers 722 papers across 372 result families, produced by an unreleased internal model trained with reinforcement learning techniques, not a model called GPT-7.
  • The problems were genuinely unsolved research questions , not benchmark tests, and the model was reportedly run against roughly 4,000 problems before researchers picked the results worth publishing.
  • Not every result is fully verified. Some papers include proofs checked by the Lean proof assistant, which mechanically verifies each logical step, but OpenAI itself says some results lack this level of checking and could contain errors.
  • The most talked-about single result, a “quasi-Riemann hypothesis” proof, is not a proof of the Riemann hypothesis itself. It proves a weaker statement about a related function, which is a real contribution but a different claim entirely.
  • An outside advisory group including respected mathematicians reviewed how the results should be communicated , but explicitly did not endorse the correctness of the mathematics, a distinction that matters a lot and gets lost in headlines.
  • Claims about counterexamples, disproofs, and “solved problems” require careful reading because a meaningful share of results reportedly involve disproving assumptions rather than proving new theorems.
  • Both Anthropic and OpenAI’s results point to the same underlying shift : giving models time to reason, use tools, and check their own steps produces output that looks a lot like real mathematical progress, even if the old “it just predicts the next token” dismissal no longer fully explains what’s happening.

How does the Lean proof checker fit into all this? #

Lean is software that verifies whether a mathematical argument, written in a very precise formal language, actually follows from its stated rules and assumptions. It doesn’t read natural language and nod along. Every step either checks out logically or it doesn’t.

This matters because it removes a huge source of uncertainty. When a chatbot produces a confident-sounding proof in plain English, readers have to trust the model’s tone and structure. When a proof is translated into Lean and the checker accepts it, you have a mechanical guarantee that the logic holds, step by step, given the starting assumptions.

The catch: formalizing a proof in Lean is a separate, often grueling task from discovering the proof in the first place, and Lean can only confirm that the formalized version is internally consistent. It can’t tell you whether the formalized statement is actually the same question the paper claims to answer. Mathematicians still have to check that framing by hand. OpenAI acknowledged that some of its released results don’t yet have Lean-checked versions, and flagged that unchecked work might contain mistakes.

Is the OpenAI math release actually a bigger deal than it sounds? #

In some ways, yes. The scale alone is unusual: 722 papers is not a cherry-picked screenshot of a chatbot nailing one hard question. The problems were drawn from a new, harder test set that OpenAI reportedly built because its older math benchmarks had become too easy to show meaningful progress between model versions. The model was tested against around 4,000 problems total, and the published papers represent the subset OpenAI judged important enough to group and release.

That framing matters because it rules out some lazy comparisons. You can’t divide papers by problems and get a clean “success rate,” since one problem can spawn several related results. You also can’t cleanly compare the “about 3 hours of computer work per result” figure OpenAI cited to the much larger effort behind its earlier Navier-Stokes project, which reportedly involved around 10,000 AI agents working for about 88 hours. Different problems, different methods, different units of measurement. Treating that as “the model got thousands of times better in weeks” misreads the numbers.

Where the release is weaker than the hype: reports that it solved a meaningful fraction of the world’s most important unsolved problems don’t appear to be backed by anything in the official materials. OpenAI has said explicitly that the list isn’t a ranking of importance. One breakdown circulating online claimed roughly a fifth of the results are disproofs (finding a single counterexample that breaks an assumed rule), but that figure is an outside estimate, not an official tally.

What does the Riemann hypothesis result actually prove? #

The Riemann hypothesis is one of the most famous unsolved problems in mathematics, concerning how prime numbers are distributed. It remains on the Clay Mathematics Institute’s list of unsolved Millennium Prize problems, alongside P vs NP.

The new paper reportedly proves something called the “quasi-Riemann hypothesis,” a weaker statement about a related mathematical function that researchers use to study primes. Proving a weaker version does not get you the stronger, full result. That distinction is easy to lose in a headline that drops the word “quasi,” turning a genuine partial step into a claim the paper never makes. The underlying contribution, the Catalan’s constant irrationality-style claim also included in the release (proving a number can’t be written as a simple fraction), can sound trivial in plain English while being genuinely hard to establish rigorously. None of this should be dismissed, but none of it should be inflated into “Riemann hypothesis solved” either.

How does this compare to Anthropic’s Claude results on Fermat’s Last Theorem? #

Anthropic has reported Claude working on problems including a Lean-formalized treatment connected to Fermat’s Last Theorem and progress on bounds related to the Riemann hypothesis. The pattern looks similar to OpenAI’s release: a frontier model given time, tools, and a formal proof checker produces results that would have sounded implausible for a chatbot a couple of years ago.

The comparison is hard to make precise because the two companies have published different amounts of supporting detail, and neither has released full technical specifics on model size, training cost, or exact methodology. What’s consistent across both is the role of formal verification. Claims checked against Lean or a similar proof assistant carry more weight than claims resting on a model’s confident prose. Claims without that backing deserve the same skepticism you’d give a human mathematician’s unchecked conjecture.

Did outside mathematicians actually endorse these results? #

Not in the way some coverage implies. OpenAI’s release involved an independent advisory group that includes well-known mathematicians, hosted at the Institute for Advanced Study. The group’s role was advising on how the results should be communicated and shared responsibly, and it explicitly stated that participation doesn’t mean endorsing the results or the process that produced them.

That’s a meaningful gap between “respected mathematicians were involved” and “respected mathematicians checked every proof and signed off.” Separately, at least one working mathematician, Daniel Litt at the University of Toronto, offered specific context on one result, noting it looked like a narrow special case of a question he’d posed, one that followed from stronger work already being developed by one of his own students. That kind of specific, grounded reaction is the useful signal here, far more than general astonishment at the volume of output.

Does any of this amount to AGI? #

No, not on its own. Artificial general intelligence typically means a system that performs well across many different kinds of tasks, not just one domain. A large batch of math papers, even verified ones, doesn’t demonstrate how a model handles unfamiliar situations outside math, whether it can operate without human oversight, or how it performs on tasks with no formal checker available. It’s fair to call these results strong evidence of real progress in mathematical reasoning. It’s a stretch to call it evidence of general intelligence.

Frequently Asked Questions #

Did an AI actually prove the Riemann hypothesis?

No. The reported result proves a weaker “quasi” version of the hypothesis for a related function, not the full Riemann hypothesis, which remains unsolved and listed on the Clay Mathematics Institute’s Millennium Prize problems.

What is Lean and why does it matter for AI math claims?

Lean is a proof assistant that mechanically verifies whether a formalized mathematical argument follows logically from its assumptions. It gives a much stronger guarantee than a chatbot’s plain-language explanation, though it can’t confirm the formalized statement matches the original question being claimed.

Is OpenAI’s internal math model GPT-7?

Other agents ship a demo. Remy ships an app. #

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

There’s no confirmation of that. OpenAI has only described it as an unreleased internal model. Naming guesses circulating online are speculation, not official information.

How many of OpenAI’s released math results are fully verified?

Not all of them. Some papers include Lean-checked proofs, but OpenAI has said some results lack this verification and may contain errors, so the collection represents a mix of confidence levels rather than uniformly confirmed mathematics.

Does this mean AI models no longer “just predict the next word”?

Next-token prediction is still part of how these models work mechanically, but it doesn’t explain how a system produces a novel argument that holds up under formal checking. The stronger framing is that giving models time to reason, use tools, and verify their own steps produces capabilities that simple description undersells.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-math-proofs-how-a…] indexed:0 read:8min 2026-10-08 · —