Claude Formalized Fermat in Lean — The Coordination Story Developers Missed Anthropic announced that its Claude model formalized Fermat's Last Theorem in Lean 4, producing 13 million lines of code and proving 29,500 theorems in 11 days, with verification by Lean's kernel and the Rust-based checker nanoda. The first multi-agent run failed due to lack of shared state, but succeeded after integrating Prove2Me, an open-source coordination platform from Columbia University's Tianyi Peng, which provided a directed acyclic graph of theorem statements. The formalization relied on community infrastructure including Mathlib and Kevin Buzzard's FLT project at Imperial College London, and experts note that machine verification does not eliminate the need for human semantic review of theorem statements. Last week, Anthropic announced Claude had formalized Fermat’s Last Theorem in Lean 4 — 13 million lines of code, 29,500 theorems proved, 11 days of wall-clock time. Most coverage landed on the math. The headline most developers missed: the first run failed. What made the second attempt work wasn’t a smarter model. It was a shared task list. What Claude Actually Did And Didn’t Do To be precise: Claude didn’t prove Fermat’s Last Theorem. Andrew Wiles did that in 1995. What Claude produced is a formalization — a translation of Wiles’s 129-page proof into Lean 4, the interactive theorem prover, so that a computer can verify every logical step line by line. The underlying mathematics is entirely Wiles’s. Claude wrote the code. That distinction matters. But what Claude wrote is still extraordinary. The 13 million lines of Lean code are five times larger than Mathlib — the community-maintained library of formalized mathematics with over 115,000 definitions and 232,000 theorems built over years. Two independent checkers verified the result: Lean’s own kernel and nanoda, a separately implemented Rust-based checker. The proof uses only Lean’s three standard axioms. Anthropic’s full write-up is on their research page https://www.anthropic.com/research/formalizing-fermats-last-theorem . The Real Story: The First Attempt Failed Anthropic’s first multi-agent run stalled. Dozens of Claude agents working in parallel had no shared view of what was proved and what wasn’t. They duplicated work, lost context, and stepped on each other’s progress. The run broke down. The fix was Prove2Me https://github.com/prove2me/prove2me workspace , an open-source platform from Columbia University’s Tianyi Peng. Prove2Me maintains a directed acyclic graph of theorem statements — a shared map of every theorem in the project, its dependencies, and its proof status. When added mid-run, agents could consult the graph to find open theorems, claim work, and record completions without collision. Each theorem is an immutable object with a separate statement file and proof file, which also sped up Lean’s compilation significantly. The engineering lesson here extends well beyond math. Parallel autonomous agents fail at scale without shared state. The pattern Prove2Me implemented — immutable work units, a coordination graph, parallel agents consuming a task queue — maps directly to any large agentic workflow. If you’re building multi-agent systems, this is the architecture worth studying. Machine-Verified Is Not the Same as Human-Understood Here’s the part that didn’t make most headlines. Lean can confirm that the encoded statements follow logically from each other. It cannot confirm that every lemma’s natural-language description actually captures the mathematics it claims to express. If a theorem was stated incorrectly — wrong name, mismatched description — and a valid chain of reasoning was built on top of it, Lean would verify the chain and flag nothing. This is what some analysts are calling the semantic review gap. Formal verification has always had this property; Claude’s run made it dramatically more visible by generating 13 million lines in 11 days. Anyone running formal-methods pipelines should assume the volume of AI-generated artifacts they’re asked to trust is about to spike. Semantic review — making sure the encoded statements mean what you think they mean — remains a human job. Claude Stood on Community Infrastructure Anthropic’s announcement was careful about attribution. Claude’s run depended on three pieces of community-built infrastructure it didn’t create: Mathlib the foundation , Kevin Buzzard’s FLT project https://github.com/ImperialCollegeLondon/FLT at Imperial College London which had been formalizing the prerequisite machinery since 2024 , and Prove2Me. Buzzard, who has been funded through September 2029 to complete the formalization by traditional means, described it as an “extraordinary autoformalization achievement” — and noted, pointedly, that it “changes nothing on the mathematics.” The open-source dependency is worth sitting with. A frontier AI lab’s landmark research run was made possible — and in fact unblocked — by open-source tools built by academic researchers. That’s not a footnote. It’s a structural fact about where frontier AI capability actually comes from. What Developers Should Do With This A few concrete things. Lean 4 https://lean-lang.org is now proven at the most extreme scale in the language’s history. If you work in formal methods, or you’ve been curious about theorem provers, the ecosystem is meaningfully stronger than it was a month ago — and the Lean-MCP bridge means proof state is queryable by LLM agents directly. Prove2Me’s coordination model is open source and worth examining as a template for parallel agentic tasks. And if you run formal verification pipelines professionally, start thinking now about how you’ll handle AI-generated proof artifacts at volume. The semantic review problem is not going away — it’s getting bigger.