Last week OpenAI uploaded 722 math manuscripts to GitHub. Each one was generated by an internal unreleased model that spent roughly three hours of ChatGPT Pro compute chewing through roughly 4,000 math problems. Here's the weird part: almost every result came from a single prompt handed to one AI agent.
I'm reading this differently than the headline-friendly framing of "OpenAI solves math problems." The fact that 722 results collapsed into a single prompt-and-agent execution is the actual story. It means the diversity in the outputs isn't coming from the model exploring different approaches or discovering independent solutions. It's one branch of reasoning, one path through the search space, repeated and pruned into 722 versions.
That changes what we're looking at. You're not seeing a model that groks multiple mathematical traditions or can solve the same problem five different ways. You're seeing a model locked onto one method, one framing, one internal representation that happens to work. Scale the generation up, ship 722 results, and it looks like breadth. In reality it's depth of a single execution.
Scott Aaronson flagged something else worth sitting with. He noted that frontier labs have begun quietly checking whether their newest math-capable models can break actual cryptographic primitives. OpenAI's October 6 release included proofs for the Unique Games Conjecture and claimed work on L=BPL. Aaronson points out that the model only solved around 5% of 8,000 attempted open problems at 3 hours of compute each. But the fact that labs are now stress-testing their models against encryption, and doing it quietly, suggests they're taking seriously the possibility that math models could become tools for breaking real security.
The proofs themselves have a tell. Mathematicians looking at the writeups describe them as "written by someone on psychedelics." That's funny on the surface. It's also a red flag about the reasoning process. The model can produce valid Lean certificates (formal proofs that check out), but the natural language explanations are incoherent. That's a model that knows when it's right without understanding why. It's pattern-matching at scale, not insight.
Here's what interests me: OpenAI published this work. Dumped it public. Framed it as a showcase. But the details, one prompt, one path, working solutions that read like fever dreams, paint a picture of something much more constrained than the narrative. The model found solutions in a very narrow way and repeated that narrow way 722 times.
If this is the state of math reasoning in late 2026, then either the breakthroughs are smaller than advertised, or the reasoning is more brittle than it looks.