TL;DR:On July 23, the 2026 Fields Medals were awarded. Wang Hong and Deng Yu — both Peking University alumni — became the first Chinese nationals to win in the same year. But this isn't a news article. It's a heist. Inside these four mathematicians' papers are four ways of thinking you can steal and use today.
Who won, which institution, how old they are — the news already covered that. Here's the snapshot:
| Winner | Institution | Core Contribution |
|---|---|---|
| Wang Hong | ||
| NYU Courant | 3D Kakeya Conjecture (100-year-old problem) | |
| Deng Yu | ||
| University of Chicago | Breakthrough on Hilbert's 6th Problem | |
| Jacob Tsimerman | ||
| University of Toronto | André-Oort Conjecture | |
| John Pardon | ||
| Stony Brook University | 3D Hilbert-Smith Conjecture |
What I care about: these four people's papers contain four mental models. Models that apply just as well to AI engineering, writing, and product decisions as they do to mathematics.
This isn't "math for beginners." This is "methodology theft."
What's the Kakeya conjecture? One sentence: take a unit line segment in 3D space, rotate it 180°. What's the minimum volume it sweeps out? Zero.
The 2D case was solved in 1917 — the area can approach zero. But 3D? Unsolved for over a century. Intuition says the dimension of a 3D Kakeya set should be 3, but there was no rigorous proof.
In 2025, Wang Hong and Joshua Zahl dropped a 127-page paper with 14 figures, rigorously proving: in three-dimensional space, any set containing unit line segments in every direction must have Minkowski and Hausdorff dimension exactly 3.
The core method, in their own words: multiscale analysis.
You don't figure out the arrangement of all tubes in 3D space at once. You split the problem into different scales. At the coarse scale: distribution patterns — which directions cluster together inside a convex set? At the fine scale: local density — if tubes are packed tight in a region, what's the volume contribution? The two scales are handled separately, then stitched together with combinatorial estimates.
Concretely, they studied this: if δ-tubes in 3D can't be packed too many into the same convex set, then their union must have nearly maximal volume. This seemingly narrow intermediate result was exactly enough to derive the full Kakeya conjecture.
The impact goes far beyond pure math. CT/MRI reconstruction algorithms rely on Fourier analysis — the exact domain of the Kakeya conjecture. MIMO antenna beamforming and radar direction estimation directly benefit from this work.
But what you should remember is the method: when facing a complex problem, go coarse first, then fine.
Don't try to figure out "how to double my blog traffic" all at once. First, categorize at the coarse scale: tool reviews get 3-10x more reads than industry analysis, and 95%+ of viral hits come from recommendations. Then fine-tune within each category: how to structure headlines with numbers, what cover style works, what time to publish. Separating scales makes the problem solvable.
Start with an analogy.
The training rule for neural networks is dead simple — gradient descent, nudging a trillion parameters by a tiny amount each step. This micro-rule is boring to the point of absurdity. But after training, the model suddenly reasons, writes code, edits articles. Stare at those trillion parameters — you won't see any of that intelligence hiding there.
This is essentially the same problem Deng Yu studies in wave turbulence theory.
In 2023, Deng and collaborator Zaher Hani published a 138-page paper with 44 figures. What they did: starting from the cubic nonlinear Schrödinger equation (describing light propagation in fiber optics, ocean waves), they rigorously derived the wave kinetic equation under a scaling limit.
The setup: inside a large box (side length L), waves interact very weakly (parameter α). As L → ∞, α → 0, with the scaling relation α ~ L⁻¹, the system's long-time statistical behavior is fully described by the wave kinetic equation — valid all the way up to the kinetic time scale T ~ α⁻².
Translation: countless simple waves, each obeying the same micro-rule → given enough time → predictable macro-level statistical laws emerge automatically.
This framework directly impacts plasma physics (laser-plasma interaction simulations for fusion), atmospheric nonlinear wave modeling in weather forecasting, and oceanography for long-time current evolution.
Together with Hani and Ma Xiao, Deng also made a breakthrough in another direction: starting from hard-sphere dynamics, they rigorously derived the Boltzmann equation on timescales far exceeding previous theorems. This advances the core of Hilbert's 6th Problem (axiomatization of physics, proposed in 1900).
For you, remember this: micro-rules × long time × many particles = predictable macro-laws. This isn't just math. Writing one article a day looks like a micro-behavior. After 100 articles, macro-patterns will emerge on their own — which headlines get traffic, which topics explode, what opening the algorithm loves.
Trust compounding, not shortcuts. This isn't motivational fluff. Deng Yu proved it with 138 pages of mathematics.
Tsimerman's André-Oort conjecture paper, v2, has a note I've reread several times:
"The previous version (v2) claiming to prove the general case has a serious error, as was kindly pointed out by Klingler, Ullmo and Yafaev. As such, the article is being reverted..."
Translation: the previous version had a serious error. Retracted.
You read that right. A Fields Medal winner's paper had a "serious error," was called out by peers, and withdrawn. Then he fixed it, resubmitted, and won mathematics' highest honor for that very work.
It wasn't just "being smart." The method matters more.
Tsimerman proved the André-Oort conjecture for A_g — a core problem about the distribution of "special points" (CM points) of elliptic curves in high-dimensional moduli spaces. Elliptic curves are the foundation of modern cryptography — the distribution of these special points directly affects the security of certain post-quantum cryptographic schemes.
But the real reason he won isn't the conclusion. It's the tools he assembled. He stitched together three fields that look completely unrelated:
Tools from any single field weren't enough. Combined, they were exactly enough.
Together with Bakker and others, he leveraged the o-minimal framework to further develop GAGA theory, solving the Griffiths conjecture about period mapping images — building a new bridge between model theory and Hodge theory.
This is exactly the state of AI right now. The next breakthrough probably isn't "better Transformers." It's grafting algebraic topology onto neural network architectures, or redescribing training dynamics in the language of wave turbulence theory.
Tsimerman teaches you two things:
Pardon proved the 3D Hilbert-Smith conjecture. 24 pages.
The question: what kinds of symmetry can "live in" three-dimensional space? Can the p-adic integers Z_p? Pardon's answer: no.
He later made even more contributions in symplectic geometry — virtual fundamental chains, localization of wrapped Fukaya categories, curve counting on Calabi-Yau threefolds (the MNOP conjecture). This work is now the mathematical foundation for topological quantum error-correcting codes — quantum computers haven't been built yet, but their error correction is already being built on his foundation.
But what fascinates me most is a single methodological phrase in his Hilbert-Smith paper: "Approach is local on M."
The entire proof doesn't need to understand all geometric properties of a 3-manifold. He only needs to find a Z_p-invariant open set U on the manifold M, locate an incompressible surface F inside U, analyze the mapping class group homomorphism induced by Z_p's action on the isotopy class of F — and the contradiction emerges.
Don't scan the whole picture. Find one window. The information inside that window is enough to derive the global conclusion.
This is directly usable.
You want to analyze data from 55 AICDragon articles to figure out what determines readership. You could run statistics on all of them — but that's making work for yourself. Just pick three: the highest-read, the lowest-read, the median. The differences between them already contain most of the patterns.
You want to debug an encoding corruption. You don't need to trace OpenClaw's entire code path. Just look at the first 4 bytes of one file — E9 94 98 3F
— it's UTF-8 BOM decoded through GBK. Root cause identified, no global investigation needed.
Pardon's lesson: it's not about having more information. It's about having the right information inside the right window.
It's not "this formula can optimize Transformers."
It's about ways of thinking.
Everything you do daily — topic selection, writing, debugging, data analysis, product decisions — has math-level complexity. The only difference is you haven't noticed.
Wang Hong says: multiscale thinking, coarse first then fine.
Deng Yu says: micro-rules × enough time = macro-laws. Trust compounding, not shortcuts.
Tsimerman says: don't grind in one field. The real breakthroughs are at the intersections. And a retraction isn't shameful — fix it and ship it.
Pardon says: don't try to understand everything. Find the right window.
These four methodologies aren't the privilege of mathematicians. They're weapons for anyone who needs to think and iterate continuously.
Including you.