# AI Just Published 722 Math Proofs. Who Checks Them?

> Source: <https://www.mindstudio.ai/blog/ai-verification-crisis-in-science/>
> Published: 2026-10-07 00:00:00+00:00

# AI Just Published 722 Math Proofs. Who Checks Them?

OpenAI released 722 AI-generated math manuscripts in one month. Here's why verifying them matters more than the results themselves.

## What happened with OpenAI’s 722 math manuscripts?

OpenAI released 722 mathematical manuscripts produced by an unreleased internal model, posted publicly on GitHub along with reasoning traces and Lean verification files. The releases span 17 subject areas, from prime number theory to plasma physics to quantum circuits, and touch on problems connected to some of the hardest open questions in mathematics, including partial progress related to the Riemann hypothesis and the Birch and Swinnerton-Dyer conjecture. None of the Millennium Prize Problems were solved outright. But the pace of output, and the fact that almost no human mathematician could independently check more than a sliver of it, is the real story.

## TL;DR

- OpenAI’s output on these math problems went from roughly 10 results in August to over 100 in September to 722 in October, a pattern of **roughly 10x growth per month** .
- Each manuscript reportedly took **about three hours of compute equivalent to ChatGPT Pro’s thinking mode** , far less than the years or decades human mathematicians typically spend on comparable problems.
- The model made **partial progress on named open problems** , such as pushing a bound related to the Riemann hypothesis further than prior human results, without resolving the underlying conjectures.
- Many of the proofs include **machine-readable Lean verification** , which helps automate some checking but does not replace the human judgment needed to confirm a result is meaningful and correctly framed.
- The real bottleneck isn’t generation, it’s **verification capacity** : there are only a few hundred to a thousand people worldwide qualified to review proofs at this frontier, and each review can take days or weeks.
- Math is being treated as a **testing ground** for what happens when frontier AI models start producing results faster than any field can check them, with biology, physics, and economics likely next.
- The bigger open question is whether these breakthroughs stay **comprehensible to humans** at all, or whether AI-driven discovery eventually outpaces our ability to follow or explain it.

## One coffee. One working app.

You bring the idea. Remy manages the project.

## Why does verification matter more than the results themselves?

Generating a proof and confirming a proof are different problems. A model can produce a document that looks like rigorous mathematics, uses correct notation, and follows plausible logical steps, without that document actually being valid. Mathematics has an advantage here that most fields don’t: formal proof checkers like Lean can mechanically verify that each logical step follows from the last, once the proof is translated into that formal language. That’s part of why OpenAI paired many of these manuscripts with Lean verification artifacts.

But machine-checkable syntax is not the same as mathematical significance. A formally valid proof can still be trivial, redundant, or solving a reformulated version of a problem that misses the point. Judging that requires a mathematician who understands the specific subfield, and for many of these 722 results, that pool of qualified reviewers is small. Some topics are specialized enough that maybe a few hundred people on earth, maybe closer to a thousand, have the background to evaluate a single paper properly. Multiply that bottleneck across 722 manuscripts released in a single batch, and you get a backlog that could take the mathematical community a very long time to clear.

## How fast is AI-generated math actually accelerating?

The growth curve is the part worth paying attention to, more than any individual result. Based on OpenAI’s own release pattern, output went from about 10 results in August to over 100 in September to 722 just six days into October. That’s consistent with roughly a 10x increase month over month. Whether or not that exact rate holds going forward, the trajectory mirrors a pattern seen repeatedly in AI capability growth: early results look unimpressive or shaky, skeptics dismiss them, and then capability compounds quickly once the underlying approach starts working.

What makes this particular case notable is that the model producing these results is not a specialized math engine. It’s framed as a general large language model, one capable of writing an email or a poem as well as advancing partial results across seven or more distinct mathematical subfields simultaneously. Human mathematicians typically specialize deeply in one or two areas because the learning curve to even approach the frontier takes years. A general model operating across many frontiers at once, with no need to specialize, is a structurally different kind of contributor to a field.

## What does this mean for other sciences like biology and physics?

Mathematics works as a proving ground for AI-driven discovery because its outputs are, in principle, checkable. A proof is either valid or it isn’t, and tools like Lean can formalize large parts of that check. Biology, physics, chemistry, and economics don’t have an equivalent formal verification layer. A claimed result in cell biology or materials science usually requires wet-lab replication, instrumentation, time, and money, not just a logical review.

## Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

That means if AI models start generating hypotheses, models, or claimed results in those fields at anywhere near the rate math output is currently growing, the verification bottleneck gets worse, not better. There’s no Lean-equivalent for most experimental sciences. The scramble to check 722 math manuscripts is a preview, on relatively favorable terrain, of a much harder problem coming to fields where confirming a result might take months of lab work rather than a close reading of a formal proof.

## Is any of this actionable yet?

None of the headline Millennium Prize Problems have been solved. What’s happened instead looks more like a tightening circle: partial results, bounds, and special cases around several of the hardest open problems in math have been pushed further than human mathematicians had gotten, without fully resolving the central conjectures. That’s meaningful progress, but it’s not yet the kind of result that changes technology or policy overnight.

Where the impact is more plausible in the near term is inside AI research itself. Mathematical breakthroughs, particularly ones touching linear algebra, optimization, and related areas, tend to feed back into how machine learning models are built, trained, and run on hardware. Other AI research efforts, including Google’s work on using AI for algorithm discovery and related projects, point toward the same loop: AI-assisted math improving the tools used to build the next generation of AI. That recursive effect, more than any single proof, may be the most consequential outcome of this release.

## Will we still understand the math AI produces?

This is the open question with no clear answer yet. Right now, a relatively small number of specialists can review and explain these results, and tools like ChatGPT can help non-experts get a rough sense of what a given proof claims. That keeps the work tethered to human understanding, even if most people will never personally verify it.

The risk is a scenario where AI systems keep producing results at compounding speed while the pool of humans capable of following the reasoning doesn’t grow at the same rate. Mathematicians have started raising this explicitly, including questions about whether maintaining a human community capable of understanding frontier math across generations should be treated as a deliberate goal, rather than an assumption. If verification can’t keep pace with generation, the field ends up with a large and growing stack of claimed results that nobody has fully checked, which undermines the reliability of the whole body of work regardless of how much of it turns out to be correct.

## Frequently Asked Questions

### What is Lean verification in mathematics?

Lean is a formal proof assistant that checks whether each step of a mathematical proof follows logically according to strict formal rules. Translating a proof into Lean allows a computer to confirm its internal logical validity, though it doesn’t by itself confirm the proof is meaningful or correctly framed relative to the original problem.

### Did OpenAI solve the Riemann hypothesis?

No. Reporting on this release describes partial progress, sometimes called results adjacent to a “quasi” Riemann hypothesis, where a numerical bound was pushed further than previous human results. The full Riemann hypothesis, which carries a Millennium Prize, remains unsolved.

### Why is math used as a test case for AI-driven discovery?

## Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Because math proofs can, in principle, be checked formally and rigorously using tools like Lean, unlike experimental fields such as biology or physics where confirming a claim often requires lab replication. That makes math a relatively controlled environment for observing what happens when AI starts generating results faster than humans can review them.

### How many people can actually verify these AI-generated proofs?

For many of the specialized subfields covered in this release, only a few hundred to roughly a thousand mathematicians worldwide may have the background needed to properly evaluate a single result, and thorough review can take days to weeks per paper.

### Will this affect fields outside mathematics soon?

The transcript’s sourcing suggests the more immediate impact is on AI research itself, since mathematical breakthroughs often feed back into machine learning techniques and hardware design. Broader effects on biology, physics, or economics are expected eventually, but no concrete timeline or result is established yet.
