OpenAI Releases Over 300 Mathematical Results Produced By An Internal Model OpenAI published a catalog of 722 mathematical manuscripts, grouped into 372 "result families" across 16 areas of mathematics, that it says were produced by an unreleased internal model, including claimed resolutions of famous conjectures and counterexamples to long-held beliefs. OpenAI said the model was posed roughly 4,000 problems with each result using about three hours of ChatGPT Pro thinking compute on average, and that the results have not been peer reviewed and some unformalized ones "could have issues." The catalog does not name the model, and OpenAI said it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release the work. AI is now generating so many mathematical results they’re now being dumped en masse, instead of being shared individually. OpenAI has published a large collection of mathematical results that it says were produced by an unreleased internal model. The catalog groups 722 manuscripts into 372 “result families” across 16 areas of mathematics. It includes claimed resolutions of some famous conjectures and a long list of counterexamples to beliefs mathematicians have held for decades. If the results hold up, this is a very large body of work. They have not been peer reviewed, though, and OpenAI itself says some of the unformalized ones “could have issues.” How the results were produced According to the company, OpenAI evaluates its models on open research problems, and it expanded those evaluations after its existing math evaluations saturated. Over the course of the exercise, the model was posed roughly 4,000 problems. Each result used about three hours of ChatGPT Pro thinking compute on average. The outputs were then aggregated into families and filtered for “an appropriate level of significance,” which produced the catalog. OpenAI says most results came from this same fixed procedure. It lists two exceptions: the work on a zero-free region for the Riemann zeta function, and the proof of the Hodge conjecture for CM abelian varieties. The writeup of the Re s 11/12 zero-free region was also edited by humans for readability. The readme does not say what was different about how those two were produced. The company says it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study when deciding how to release the work. On verification, it says not all results come with Lean formalizations and that it will add them as they are completed. Corrections will be recorded as new versions, and earlier versions will stay accessible. OpenAI is also releasing abridged summaries of the model’s reasoning for ten of the results, and says it is exploring community-hosted repositories for the materials. The catalog doesn’t name the model. OpenAI’s most recent public math claims came from internal models too. In May it said an internal model had disproved Erdős’s 1946 planar unit distance conjecture. Last month it shared a claimed solution to the Navier–Stokes problem https://officechai.com/ai/openai-shares-solution-to-navier-stokes-problem-created-by-a-model-significantly-more-capable-than-astra/ from a model it described as more capable than its public one, a claim that came with a dispute over credit and had not yet been examined by outside mathematicians. What’s in the catalog | Area | Families | A few of the claims | |---|---|---| | Number theory | 31 | Quasi-Riemann hypothesis zero-free for Re s 7/8 ; Hilbert’s tenth problem over ℚ; irrationality exponent of π is 2; Goldfeld’s conjecture; modularity over imaginary quadratic fields | | Algebraic and complex geometry | 36 | Hodge conjecture for CM abelian varieties and products of K3 surfaces; Fujita’s freeness conjecture; Kobayashi’s canonical-ampleness conjecture; generalized Mukai conjecture | | Real and complex analysis | 16 | Kakeya in 3D maximal and 4D Hausdorff ; Falconer distance conjecture; 3D Bochner–Riesz | | Convex and metric geometry | 15 | Mahler conjectures; logarithmic Brunn–Minkowski; optimal covering density Θ n log n | | Theoretical computer science | 40 | Unique Games Conjecture; L = RL = BPL; matrix multiplication exponent ≤ 9/4; integer multiplication below n log n | | Dynamical systems and ergodic theory | 12 | Uniform limit-cycle bound in Hilbert’s 16th problem; Rokhlin’s multiple-mixing problem; Sinai’s conjecture for the standard map | | Combinatorics | 37 | The plane is not 5-colorable; Borsuk’s conjecture fails in dimension nine; Hadwiger and Sidorenko counterexamples; Harary–Hill crossing-number conjecture | | Algebra | 18 | Serre’s intersection multiplicity positivity; Lech’s conjecture; Kaplansky zero-divisor counterexample; blockwise Alperin weight conjecture | | Probability and statistical mechanics | 29 | p c < p u for nonamenable graphs; Sherrington–Kirkpatrick results; random k-SAT thresholds; BKT scaling for the planar XY model | | Mathematical logic | 6 | Shelah’s eventual categoricity; rigidity of the Turing degrees; partition principle does not imply choice | | Group theory | 14 | Cannon’s conjecture; Thompson’s group F is nonamenable; Boone–Higman; an infinite finitely presented simple amenable group | | Mathematical physics | 25 | Anderson localization and delocalization; spacetime Penrose inequalities; the spin-one Haldane gap; exactly three mutually unbiased bases in dimension six | | Operator algebras | 19 | All nonabelian free group factors are isomorphic; Kadison’s similarity conjecture; Connes’ bicentralizer conjecture | | Topology | 18 | Hilbert–Smith conjecture; purely cosmetic surgery conjecture; Quillen’s conjecture in rational homology | | Functional analysis | 11 | Tingley’s problem; complete Crouzeix conjecture; Kirk’s fixed-point problem | | Differential geometry | 29 | Yau’s uniformization conjecture; Katok’s entropy rigidity; metric Blaschke conjecture; Donaldson’s tamed-to-compatible conjecture | | Partial differential equations | 16 | Global smoothness for relativistic Vlasov–Maxwell; De Giorgi’s conjecture in dimension eight; hot spots for simply connected planar domains | A rough count suggests several dozen of the families are counterexamples or disproofs rather than proofs of conjectures. These include the failure of Borsuk’s conjecture in dimension nine, a torsion-free group algebra with zero divisors, the Hadwiger counterexample, and a smooth surface metric with no local isometric immersion into three-space. How significant is this? The sheer breadth stands out. Several of the claims are central problems in their fields. The Unique Games Conjecture has shaped the theory of approximation hardness for nearly twenty years. Hilbert’s tenth problem over the rationals has been open for decades. The isomorphism of free group factors has been a defining question in operator algebras, and the amenability of Thompson’s group F is one of group theory’s best-known open questions. Some numeric claims would also be big moves on their own: a matrix multiplication exponent of 9/4 the long-standing public bounds are around 2.37 , an irrationality exponent of 2 for π the earlier upper bound was around 7 , and a counterexample to Borsuk’s conjecture in dimension nine the smallest dimension known before was in the sixties . The contrast with how recently AI math was being measured is large. A little over a year ago, an OpenAI model’s gold-medal-level performance at the International Mathematical Olympiad https://officechai.com/ai/an-openai-model-has-delivered-a-gold-medal-performance-at-the-math-olympiad/ was the headline milestone. The company now says its models are producing research-level results at scale. Several results, however, might b e narrower than their names suggest. The “quasi-Riemann hypothesis” result is a zero-free half-plane for Re s 7/8, not the Riemann hypothesis. The Hodge conjecture result covers CM abelian varieties and products of K3 surfaces, not Hodge in general. The Birch–Swinnerton-Dyer result applies when a Selmer group has corank zero or one. Kakeya is resolved in three and four dimensions, not in general. Almost every entry carries these qualifiers in its fine print, and they matter a lot for judging the real impact. Verification is the open question. The collection spans nearly every area of pure mathematics, and no group of referees can check 722 manuscripts quickly. Lean formalization covers “many, but not all” of them, which means some results come with machine-checked proofs and others are only as trustworthy as the first readers who get to them. Terence Tao has warned that AI-generated proofs can look superficially flawless while hiding subtle mistakes https://officechai.com/ai/ai-can-generate-math-proofs-that-look-flawless-but-makes-subtle-mistakes-that-humans-wouldnt-terrance-tao/ . The mathematicians who will do the checking are also not uniformly enthusiastic about how AI labs present the field. Tao has said OpenAI used only upbeat snippets of a much longer conversation https://officechai.com/ai/openai-took-optimistic-snippets-around-ai-of-my-conversation-for-its-ad-says-terence-tao/ with him in one of its ads. Not every result is equally deep. Open problems vary enormously in difficulty, and a catalog of 372 families almost certainly mixes landmark results with obscure ones. The descriptions don’t make it possible to tell from outside which is which. One entry, on random k-SAT thresholds, notes that another researcher, Gaia Carenini, independently resolved the threshold-existence question in a concurrent preprint made public on October 5, and credits her with priority. That is a small example of a problem that will likely get bigger: when many labs and individuals attack the same problems with similar tools, questions of who solved what first become harder https://officechai.com/ai/openai-says-it-didnt-see-any-of-buckmaster-and-alpoges-work-while-resolving-navier-stokes/ . Why it matters for the industry For AI labs, open math problems have become the next benchmark. Once competition-style questions saturated, the evaluation shifted to problems no one knows the answer to, and the output of that evaluation doubles as a credibility signal. Math is also a domain where a proof is either correct or it isn’t, which is why it’s an appealing place to show reasoning ability. The economics are worth noting. OpenAI puts the average at about three hours of Pro compute per result. Its Navier–Stokes effort was a much larger, more resource-heavy push. The talent market is moving too: Fields Medalist Jacob Tsimerman is taking leave from Toronto to join OpenAI https://officechai.com/ai/fields-medal-winner-jacob-tsimerman-joins-openai-says-ai-will-soon-do-everything-mathematicians-do-better-and-faster/ , saying AI will soon do much of what mathematicians do. What happens next will be the real test. Watch for Lean formalizations being added, for independent expert write-ups on the biggest claims, and for whether any of the manuscripts are accepted by journals. A few retractions wouldn’t undo the collection’s importance, but a long list of confirmed results would change how mathematics is done. The next few months will show which of the two this release turns out to be.