cd /news/artificial-intelligence/openai-s-722-math-proofs-all-came-fr… · home › topics › artificial-intelligence › article
[ARTICLE · art-148129] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

OpenAI's 722 Math Proofs All Came From One Prompt

OpenAI uploaded 722 machine-generated math manuscripts to GitHub, all produced by an unreleased internal model that ran roughly three hours of ChatGPT Pro compute per problem across about 4,000 problems. The results reportedly trace back to a single prompt and a single agent execution, meaning the apparent breadth reflects one reasoning path repeated rather than independent solutions. Mathematicians describe the natural-language writeups as incoherent even where the accompanying Lean certificates check out, and Scott Aaronson noted frontier labs are quietly testing math-capable models against cryptographic primitives, with the model solving only about 5% of 8,000 attempted open problems.

by read2 min views5 publishedOct 9, 2026

Last week OpenAI uploaded 722 math manuscripts to GitHub. Each one was generated by an internal unreleased model that spent roughly three hours of ChatGPT Pro compute chewing through roughly 4,000 math problems. Here's the weird part: almost every result came from a single prompt handed to one AI agent.

I'm reading this differently than the headline-friendly framing of "OpenAI solves math problems." The fact that 722 results collapsed into a single prompt-and-agent execution is the actual story. It means the diversity in the outputs isn't coming from the model exploring different approaches or discovering independent solutions. It's one branch of reasoning, one path through the search space, repeated and pruned into 722 versions.

That changes what we're looking at. You're not seeing a model that groks multiple mathematical traditions or can solve the same problem five different ways. You're seeing a model locked onto one method, one framing, one internal representation that happens to work. Scale the generation up, ship 722 results, and it looks like breadth. In reality it's depth of a single execution.

Scott Aaronson flagged something else worth sitting with. He noted that frontier labs have begun quietly checking whether their newest math-capable models can break actual cryptographic primitives. OpenAI's October 6 release included proofs for the Unique Games Conjecture and claimed work on L=BPL. Aaronson points out that the model only solved around 5% of 8,000 attempted open problems at 3 hours of compute each. But the fact that labs are now stress-testing their models against encryption, and doing it quietly, suggests they're taking seriously the possibility that math models could become tools for breaking real security.

The proofs themselves have a tell. Mathematicians looking at the writeups describe them as "written by someone on psychedelics." That's funny on the surface. It's also a red flag about the reasoning process. The model can produce valid Lean certificates (formal proofs that check out), but the natural language explanations are incoherent. That's a model that knows when it's right without understanding why. It's pattern-matching at scale, not insight.

Here's what interests me: OpenAI published this work. Dumped it public. Framed it as a showcase. But the details, one prompt, one path, working solutions that read like fever dreams, paint a picture of something much more constrained than the narrative. The model found solutions in a very narrow way and repeated that narrow way 722 times.

If this is the state of math reasoning in late 2026, then either the breakthroughs are smaller than advertised, or the reasoning is more brittle than it looks.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-s-722-math-pr…] indexed:0 read:2min 2026-10-09 · —