{"slug": "show-hn-reducing-an-openai-proof-by-25-in-a-few-hours", "title": "Show HN: Reducing an OpenAI proof by 25% in a few hours", "summary": "An AI-driven proof-simplification loop reduced OpenAI's Unique Games proof in Lean from 151,287 to 113,990 lines, a 24.65% cut, according to a Show HN writeup by the project's author. The work, run over a few hours on the openai/math repository, used Codex agents (openai-codex/gpt-6.1-sol and openai-codex/gpt-6-astra) and one Claude agent (claude-agent/claude-fable-5-1) across rounds 18–23, with the final round deploying 24 Codex agents each on a separate working copy. The final proof built from empty project build folders, passed Lean's kernel check with the same theorem statement, and used only the allowed axioms propext, Classical.choice, and Quot.sound, with no sorryAx.", "body_md": "OpenAI dropped a set of proofs, and a lot of people are trying to formulate their interpretation on what this means for their field in general.\n\nThe goal here is to add a data point to the discussion.\n\n**151,287 → 113,990 lines of Lean code: a 24.65% reduction.**\nBlank lines and comments are not counted. The theorem statement stayed the same.\n\nThis experiment asks how much AI models can shrink an existing proof. The process\nis simple: make changes, check them, save the working version, and repeat.\nThe starting point was the Unique Games proof from\n[openai/math](https://github.com/openai/math).\n\nThis was a side project over a few hours. A few prompts set up the loop, then the agents worked through repeated rounds of changes and checks.\n\nEach row uses the same counting method. It counts the proof and the local files it needs.\n\n| Stage | Models and setup | Lines left | \n|---|---|---|\n| Original | Original proof | 151,287 | \n| Rounds 18–19 | Codex agents, one after another | 141,443 | \n| Round 20 | One Claude agent | 120,980 | \n| Round 21 | Four Codex agents working at the same time | 120,571 | \n| Round 22 | 24 Codex agents, each with a separate working copy | 117,213 | \n| Round 23 | Another group of 24 Codex agents | **113,990** | \n\nModel names reported by Pi:\n\n- Rounds 18–19 and 21: `openai-codex/gpt-6.1-sol`\n- Round 20: `claude-agent/claude-fable-5-1`\n- Rounds 22–23: `openai-codex/gpt-6-astra`\n\nAll used the high thinking setting. Earlier attempts did not produce changes that were kept. The prompts asked for 10% or 40% reductions, but these were goals, not the results of each round.\n\nThe agents removed repeated code and reused existing proofs. In later rounds, they shared build results but kept their working files separate. Up to four build commands could run at once. After each group finished, their changes were combined and checked together.\n\n- Both versions built successfully from empty project build folders. Existing builds of external libraries were reused.\n- The final theorem statement matched the original. Lean's kernel, which checks proofs, checked the final proof again and accepted it.\n- The proof uses only the allowed axioms: `propext` ,`Classical.choice` , and`Quot.sound` . It does not use`sorryAx` , which allows unfinished proofs.\n- The protected model and test files, Lean 4.34.1, and the versions of the external libraries stayed the same. The extra Nanoda check was not run.\n- All new helper code is counted. Removing unrelated projects from the original\nrepository does **not** count as a reduction.\n\nThe Git history starts with `original-minimal`, followed by `simplified`.\nThe final count includes all project proof files, even any files no longer used.\n\n```\npython3 scripts/audit_proof.py --before-repo . --before-rev original-minimal \\\n  --after-repo . --after-rev simplified --output /tmp/wiggums-audit\n```\n\nSee [how to repeat the checks and read the results](https://github.com/offline-ant/wiggums-proof-loop/blob/main/REPRODUCE.md).\n\nLicense: Apache-2.0. See [source and change details](https://github.com/offline-ant/wiggums-proof-loop/blob/main/MODIFICATIONS.md).\nThis README was written by the AI that led the work, at Roelof's request.", "url": "https://wpnews.pro/news/show-hn-reducing-an-openai-proof-by-25-in-a-few-hours", "canonical_source": "https://github.com/offline-ant/wiggums-proof-loop", "published_at": "2026-10-09 20:30:07+00:00", "updated_at": "2026-10-09 20:52:04.575358+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "large-language-models", "developer-tools"], "entities": ["OpenAI", "Lean", "Codex", "openai-codex/gpt-6.1-sol", "openai-codex/gpt-6-astra", "claude-agent/claude-fable-5-1", "openai/math", "Roelof"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-reducing-an-openai-proof-by-25-in-a-few-hours", "markdown": "https://wpnews.pro/news/show-hn-reducing-an-openai-proof-by-25-in-a-few-hours.md", "text": "https://wpnews.pro/news/show-hn-reducing-an-openai-proof-by-25-in-a-few-hours.txt", "jsonld": "https://wpnews.pro/news/show-hn-reducing-an-openai-proof-by-25-in-a-few-hours.jsonld"}}