An AI System Claims a 90-Year-Old Math Prize Problem. The Real Story Is Messier OpenAI announced that an internal AI system solved the Navier-Stokes existence and smoothness problem, a Millennium Prize problem, using a swarm of about 10,000 concurrent agents that burned 130 billion output tokens and cost an estimated $15 million. The result, which claims to find a finite-time blow-up for the equations, was verified in Lean. However, the announcement is overshadowed by allegations from mathematician Tristan Buckmaster that OpenAI may have used his unpublished work without permission and attempted to manipulate credit, which OpenAI denies. On September 8, OpenAI announced that an internal AI system produced a solution to the Navier-Stokes existence and smoothness problem. If you are not deep into mathematics, that sentence means nothing. So let me unpack what happened, because the announcement is the least interesting part of this story. The Navier-Stokes equations describe how fluids move. Air over a wing, water in a pipe, blood through arteries. Engineers use them every day. The equations work. Planes fly. But there is a question mathematicians have failed to answer for about 90 years: do these equations always behave? Specifically, if a fluid starts out smooth, can the equations ever produce a solution where velocity spikes to infinity in finite time? Mathematicians call this a "blow-up" or a singularity. A real fluid cannot move infinitely fast, so a blow-up would mean the equations stop describing reality at that moment. The model itself would be broken. In 2000, the Clay Mathematics Institute put $1 million on answering this. It is one of the seven Millennium Prize Problems. Six remain unsolved. That is the company this problem keeps. OpenAI's claim: their system found a blow-up. A vortex that spirals inward, stretches like spaghetti, shrinks and spins faster, with total energy staying finite the whole time. The proof ships with a formalization in Lean, a proof-checking programming language. If the Lean check is complete and correct, the math logic is machine-verified. That is a strong claim, not just a vibe. This is the part that should grab anyone who works with AI. OpenAI heard a rumor on September 1 that two Millennium Problems had fallen. They launched an effort the same day, pointing a coordinating swarm of agents at every open Millennium Problem. The group that produced the Navier-Stokes result involved roughly 10,000 concurrent agents running on an unreleased internal model, one they describe as significantly more capable than GPT-6 Astra. The agents sent 2.7 million messages and burned about 130 billion output tokens. The solution landed on September 5, about 88 hours after launch. Lean verification took another 17 hours. Total compute cost, per OpenAI's own estimate at their press conference: around $15 million. As a side quest, about 100 agents spent 50 hours resolving the Euler regularity problem, the same blow-up question for the frictionless cousin of Navier-Stokes. Two Millennium-level results in one week. Hours before OpenAI's announcement, two mathematicians published their own results: Tristan Buckmaster of NYU and Levent Alpoge of Anthropic. They proved blow-up for the Euler equations under smooth forcing, along with two related equations, verified in Lean. That is the stepping-stone result. Their work builds on years of groundwork by Diego Cordoba and Luis Martinez-Zoroa, who developed the "forcing" technique both teams ultimately used. Then it got ugly. Buckmaster published a statement alleging that rumors of his work reached OpenAI before OpenAI started its own search. He says he was using OpenAI's Codex tools throughout his project, assumed his drafts were private, and later could not get a straight answer on whether that data touched OpenAI's model training. He alleges that in calls with OpenAI's Sebastien Bubeck, the story of "a model solved it with little human input" shifted multiple times, and that OpenAI proposed announcement deals that would have dropped Alpoge from credit, citing his Anthropic employment. He says when he threatened to go public, the reply was: "Why would you ruin your career?" Bubeck called the allegations "false and inflammatory." OpenAI's version, from their post: they never saw Buckmaster's work, their Euler proof differs from his, and they offered a joint announcement that acknowledged his priority. Noam Brown and Anthropic's Sholto Douglas then traded posts, with Douglas mockingly mirroring Brown's phrasing back at him. Buckmaster himself says he is not accusing anyone of anything and that the bigger story is what frontier models can now do. Both things can be true: the math may hold, and the credit fight around it may be a mess. Watch this space. Both sides promised fuller accounts. Even setting the drama aside, there is a real technical caveat. The Clay problem's official statement includes a forcing term, an external force applied to the fluid. That term is exactly what both solutions exploit. Many experts consider the "real" question to be the version without external forcing. Scientific American reported that some in the community may argue the problem was solved through a loophole in the official wording. Whether the prize gets awarded will take years. Clay's president says the evaluation will be "deliberately unhurried." Terence Tao raised a deeper concern before the announcement. A proof that lands as a viral announcement, with the discovery process hidden inside a corporate system, gives mathematics the answer but not the understanding. The journey of attacking an open problem, iterating on failed approaches, is where a field actually learns. If an AI swarm produces a 100-page proof no human can digest, and the process stays secret, we get the trophy and lose the lesson. Tao called this scenario a net negative for the progress of mathematics. Buckmaster's own words about the AI-assisted writeup he had to decipher: he called the first model-generated proof he received the worst thing he had ever read. Strip the scandal and three things remain. First, the frontier moved. A model that beats GPT-6 Astra "significantly," still in training, just produced a formalized solution to a problem that resisted humans for 90 years. Whatever ships next from OpenAI inherits this capability. The gap between research demos and what you use through an API is shrinking. Second, the shape of the work changed. This was not one genius model. It was 10,000 coordinated agents, group communication, cross-pollination of intermediate results, and machine-checked verification at the end. That architecture is reproducible in miniature. Multi-agent systems with verification loops are where serious results are coming from, not single prompts. Third, verification is the trust boundary. Nobody has to take OpenAI's word for the math because Lean checked it. The same pattern applies to AI-written code. The teams getting real value from AI are the ones with strong tests, type systems, and review gates. Trust the verifier, not the generator. A machine just walked through a door mathematicians spent a century knocking on. Whether that door leads somewhere worth going depends less on the model and more on whether the humans fighting over the credit remember to let the rest of us understand what happened. Sources: OpenAI's announcement, statements from Tristan Buckmaster, coverage from New Scientist, Scientific American, and Terence Tao's commentary on Mathstodon. tags: ai, mathematics, machinelearning, news