Tao at ICM: AI Makes Proofs Cheap and Trust Expensive Terence Tao, in his public lecture at the International Congress of Mathematicians on July 24, argued that AI will soon generate proofs faster than humans can verify, digest, or trust, creating a crisis of 'proof indigestion.' Tao cited the First Proof challenge, where AI solved seven of ten research-level problems at publication quality in May 2026, up from two of ten in February, at costs of $10 to $1,000 per problem. He warned that mathematics institutions built for proof scarcity are unprepared for an era of proof abundance, with AI-generated submissions already piling up unverified in databases like Erdős Problems. AI https://sourcefeed.dev/c/ai Article Tao at ICM: AI Makes Proofs Cheap and Trust Expensive The ICM 2026 lecture reframes AI-for-math as a pipeline problem software teams will recognize immediately. Rachel Goldstein https://sourcefeed.dev/u/rachel goldstein Terence Tao gave the public lecture at the International Congress of Mathematicians in Philadelphia on July 24, and he opened with a move most AI commentators can't manage: he refused to argue about whether AI can do mathematics. His slides https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf frame the capability question as a "family of conjectures" — at some point, some AI tools, at some cost, with some supervision, will do some research-level math at some quality — then set the whole dispute aside. Assume a reasonably strong version is true, he says, and ask the question nobody's asking: what happens to everything downstream of the proof? That's the right question, and developers already know the answer, because we've spent the last three years living it. Generation is no longer the bottleneck Tao's core argument is a pipeline diagram. A theorem doesn't count when it's proved. It has to be verified, then written up so humans can follow it, then peer-reviewed and accepted, then — slowest and most valuable of all — digested and canonicalized into the textbooks and shared theory of the field. AI is spectacular at accelerating exactly one stage of that pipeline: generation. Every stage after it is still throttled by scarce human attention. His prediction is what he calls "proof indigestion": AI-generated proofs piling up unverified, verified proofs waiting on readable write-ups, correct and well-written papers overwhelming peer review. Mathematics, he argues, is transitioning from an era of proof scarcity to one of proof abundance, and its institutions were built for scarcity. The early evidence is already visible. Erdős Problems https://www.erdosproblems.com , the community database of Paul Erdős's open questions, now holds dozens of AI-generated proof submissions that no human expert has volunteered to check — in several cases the submitters declared themselves unqualified to vouch for their own uploads. One Erdős problem open for 60 years reportedly fell to a ChatGPT session in about 80 minutes. Tao's uncomfortable question: could we end up with a verified proof of a major result that no human understands well enough to explain? What the First Proof numbers actually say The one capability data point Tao does endorse is the First Proof https://1stproof.org challenge, and it's worth reading carefully because it cuts both ways. The first batch, in February, was chaos in miniature. Ten unpublished research-level problems; OpenAI claimed six of its ten solutions were likely correct; refereeing by the organizing mathematicians confirmed only two, and one of those closely matched an existing proof. The gap between claimed and confirmed is precisely the "reporting bias and non-scientific incentives" Tao warns pollute nearly all public evidence about AI-for-math. So for the second batch, tested May 28 under controlled conditions the organizers ran themselves now as a registered nonprofit , everything was refereed by experts for correctness and exposition. Result: seven of ten problems solved at publication-level quality by at least one of four AI harnesses, at compute costs of $10 to $1,000 per problem. From two-of-ten confirmed to seven-of-ten refereed in three months. Whatever your prior on the capability conjecture, the direction of travel is not ambiguous — which is exactly why Tao spends his hour on the downstream consequences instead. Software ran this experiment first Here's the synthesis Tao gestures at but doesn't spell out: mathematics is about to hit the wall software teams hit in 2023–2024. AI code generation made pull requests nearly free, and every engineering org discovered that review capacity, not authoring capacity, was the real constraint. Maintainers of major open-source projects started publishing AI-contribution policies after drowning in plausible-looking slop reports. The mathematicians' version is arriving on schedule, and their responses map almost one-to-one onto ours. The Leiden Declaration https://leidendeclaration.ai on AI and Mathematics — published June 2 and endorsed by the International Mathematical Union — is essentially a CONTRIBUTING.md for the discipline: disclose tool use, make your work easy to review, and keep correctness responsibility with the human authors regardless of what generated the text. Tao's own rule of thumb is sharper: if the authors can't give a clear, expert-level talk on their result, it shouldn't be published. That's "don't merge a PR you can't explain," elevated to publication policy. Tao also lands a point that software culture hasn't fully internalized: exposition can be over -optimized. A human-written proof retains "natural friction" at the hard parts — the places where the author struggled signal where the reader should slow down. An AI-polished proof sands all of that away, presenting routine and difficult steps as equally digestible, which reads beautifully and teaches nothing. Anyone who has reviewed a flawless-looking AI-generated diff that hid its one load-bearing subtlety in uniform, confident prose knows exactly what he means. Where the leverage is If you build developer tools, the actionable reading is that the opportunity has moved down the pipeline. The generation side is a frontier-lab arms race. The digestion side — verification infrastructure, review tooling, provenance — is wide open, and Tao names the stack explicitly: proof assistants like Lean https://lean-lang.org , Rocq, and HOL, plus the Mathlib library that gives formalization its standard-library backbone. Formal verification is the one downstream stage that machines can genuinely accelerate; it's CI for theorems, and it's the reason the math community may weather proof abundance better than, say, scientific publishing at large, which has no compiler for truth. The honest caveat: Tao's working hypothesis is still a hypothesis, and he's careful to say the evidence remains uncontrolled and incentive-soaked. First Proof is one benchmark, run twice. But the strength of this lecture is that it doesn't depend on the timeline. Goodhart's law — optimize the metric of "solved problems" and you'll flood the zone with unverifiable solutions — bites whether strong capability arrives in two years or ten. Mathematics is getting its values audit before the crisis fully lands. Software, which got the crisis first and the values audit piecemeal, could stand to read the slides. One more thing, from a footnote on slide five: "All em-dashes in these slides were human-generated." The man knows his audience. Sources & further reading - Terence Tao: Mathematics in the Age of AI pdf https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf — teorth.github.io - First Proof is AI's toughest math test yet. The results are mixed https://www.scientificamerican.com/article/first-proof-is-ais-toughest-math-test-yet-the-results-are-mixed/ — scientificamerican.com - First Proof's second batch of math problems test AI https://www.math.harvard.edu/first-proofs-second-batch-of-math-problems-test-ai/ — math.harvard.edu - First Proof Second Batch https://arxiv.org/abs/2606.18119 — arxiv.org - Leiden Declaration on Artificial Intelligence and Mathematics https://leidendeclaration.ai — leidendeclaration.ai - AI Will Be Top of Mind at ICM, Math's Biggest Conference https://www.simonsfoundation.org/2026/05/04/ai-will-be-top-of-mind-at-icm-maths-biggest-conference/ — simonsfoundation.org Rachel Goldstein https://sourcefeed.dev/u/rachel goldstein · Dev Tools Editor Rachel has been embedded in the developer tooling ecosystem for nearly eight years, covering everything from IDE wars and package-manager drama to the quiet rise of AI-assisted coding. She has a soft spot for open-source maintainers and an unhealthy number of terminal emulators installed on a single laptop. Discussion 0 No comments yet Be the first to weigh in.