{"slug": "deepseek-wins-imo-gold-on-12-cents", "title": "DeepSeek wins IMO Gold on 12 cents", "summary": "DeepSeek V4 Flash scored 30/42 on IMO 2026 problems in Cline, clearing the 29-point gold medal cutoff for just $0.12, roughly 140 times cheaper than Claude Fable 5. The benchmark, run by Cline, tested eight models including GPT-5.6 Sol, which scored a perfect 42/42, and used blind grading by GPT-5.5 and Claude Opus 5 with Gemini 3.1 Pro resolving disagreements.", "body_md": "[Insights](https://cline.ghost.io/tag/insights/)\n\n# DeepSeek wins IMO Gold on 12 cents\n\nBenchmarking eight AI models on IMO 2026 problems in Cline, we found DeepSeek V4 Flash scored 30/42, clearing the gold medal cutoff for just $0.12. See how open-weight models compare with GPT-5.6 Sol, Claude Fable 5, Kimi K3, and more.\n\nWe wanted to know the cheapest possible way to win an IMO gold medal, so we ran eight models against IMO 2026 problems in Cline and had the proofs graded blind.\n\nThe answer was **12 cents**.\n\nFrontier models scoring perfect runs has [been covered extensively](https://github.com/deedy/imo-2026?ref=cline.ghost.io) . Our interesting find is that open-weight models now score gold too. Our DeepSeek V4 Flash run hit 30/42, clearing the 29-point gold cutoff. This was roughly 140 times cheaper than Claude Fable 5. In this post, we explore the problem setup and judging, and share all the traces from these runs.\n\n## Problem Setup\n\nThe original benchmark was an 8 × 6 matrix: eight models and six IMO 2026 problems, making a total of 48 effective cells.\n\n- GPT-5.6 Sol\n- Claude Fable 5\n- Kimi K3\n- DeepSeek V4 Flash\n- Qwen 3.6 35B A3B\n- DeepSeek V4 Pro\n- GLM 5.2\n- MiMo V2.5 Pro\n\nEach model received the problem statement and one instruction to submit its strongest complete final solution. A submission tool was used to finish the run.\n\nAfter the original symmetric panel revealed some harness bugs, we fixed the tool-calling issues within the Cline harness to give every model its best shot at scoring, so that any lost points reflect actual model performance rather than tool-calling failures.\n\n### Judging\n\nThe proofs were graded anonymously on the IMO 0–7 scale. Two independent, model-blind graders, GPT‑5.5 and Claude Opus 5, scored each submission, and disagreements were resolved by Gemini 3.1 Pro. GPT-5.6 Sol remained the reference perfect scorer at 42/42. Other groups have independently reproduced that perfect score [here](https://www.linkedin.com/posts/akashnil-dutta-72894a100_imo-2026-performance-by-chatgpt-56-pro-activity-7483603460488273922-g1Yh/?ref=cline.ghost.io), [here](https://github.com/deedy/imo-2026?ref=cline.ghost.io), and [here](https://www.linkedin.com/posts/eugene-nazirov_the-headline-is-the-4242-imo-score-the-activity-7487441869992267776-OmY9/?ref=cline.ghost.io) (although their harness settings differ from ours). The gold medal score cutoff this [year was 29](https://www.imo-official.org/results/individual/year/2026/?ref=cline.ghost.io).\n\nEach candidate proof was placed into an immutable anonymous batch. The graders saw the problem, proof, rubric, and comparison reference (including Lean solutions), but not the candidate model's identity, to avoid bias. They checked that the IMO scoring rules were followed and that the final results passed validation.\n\nReward-hacking guardrails included disabling internet use in the system prompt and exposing only the submit-solution tool. No prior context about the problems like Lean artifacts, rationales, etc. was provided, so each model had to rely on its own ability. We also checked traces to make sure no browser use was done. And since the IMO 2026 problems are new, they aren't in these models' training sets yet.\n\nOne caveat is that IMO problems, unlike a traditional coding benchmark, are not a stable population estimate: small changes like prompt tweaks, retries, and provider issues (for open-weight models) can shift the scores a bit. We ran multiple attempts for all models and picked the best score on each problem rather than using standard measures like Pass@k. Note that the prices in the table below reflect the single best run for each model, not the net cost of all the exploration needed to work through tool-calling issues.\n\nIn general, variance of results is always something to keep in mind with any benchmark. We used an LLM as a judge instead of unit tests (math olympiad problems often don’t come with unit tests) or a formal Lean grader. You can look at unofficial Lean solutions [here](https://github.com/AxiomMath/IMO2026?ref=cline.ghost.io).\n\nAs promised, we have attached all the solution traces from the different models [here](https://gist.github.com/arafatkatze/fc08975b473205e52f272d4e7b2ad4b1?ref=cline.ghost.io); these traces include prices, generation costs, and the solutions offered by each model.\n\n## Model Scores and Prices\n\n| Model | Best score | P1–P6 | Price |\n|---|---|---|---|\n| GPT-5.6 Sol | 42/42 |\n7 · 7 · 7 · 7 · 7 · 7 | $3.2336 |\n| Claude Fable 5 | 41/42 |\n7 · 7 · 6 · 7 · 7 · 7 | $17.1956 |\n| Kimi K3 | 35/42 |\n7 · 4 · 7 · 7 · 7 · 3 | $5.1328 |\n| DeepSeek V4 Flash | 30/42 |\n7 · 4 · 3 · 7 · 7 · 2 | $0.1215 |\n| DeepSeek V4 Pro | 30/42 |\n7 · 4 · 2 · 7 · 7 · 3 | $0.4761 |\n| MiMo V2.5 Pro | 30/42 |\n7 · 2 · 3 · 7 · 7 · 4 | $1.0650 |\n| GLM 5.2 | 21/42 |\n7 · 0 · 0 · 7 · 7 · 0 | $2.3855 |\n| Qwen 3.6 35B A3B | 16/42 |\n7 · 2 · 1 · 2 · 3 · 1 | $0.2251 |\n| IMO 2026\n|\n\n**16/42**### History of AI models in IMO\n\nOriginally, in 2024, DeepMind's custom AlphaProof models (which were not released to the public) got an IMO silver, and back then the problems were converted to Lean. The models solved one problem within minutes and took up to three days to solve the others; note that humans solved the problems over two days, in sessions of 4.5 hours each.\n\nIn 2025, an advanced version of Gemini Deep Think achieved a gold medal score working directly from the official natural-language problem descriptions.\n\nIn 2026, we have made great progress in IMO and other math olympiad problem solving. For starters, we have migrated to much better mathematical benchmarks like [Riemann bench](https://arxiv.org/abs/2604.06802?ref=cline.ghost.io), which are far more intricate and long-horizon, and models are now able to solve many [unsolved Erdos problems.](https://openai.com/index/model-disproves-discrete-geometry-conjecture/?ref=cline.ghost.io) There are many other interesting solutions to math problems like [Hadamard matrices](https://x.com/__alpoge__/status/2087504785952182273?ref=cline.ghost.io) and the [Jacobian conjecture](https://x.com/__alpoge__/status/2079028340955197566?ref=cline.ghost.io) that were released this year, but there’s something very special that came from the open-weights community.\n\n### What's unique about DeepSeek V4 Flash\n\nThe most important part of the progress this year has been an order-of-magnitude improvement in the performance of open-weight models. Notably, DeepSeek V4 Flash isn’t post-trained specifically for Olympiad problems we hit the model endpoint directly and still got a gold medal. It has 284B total MoE parameters, with about 13B activated per token. While it's not exactly a small model, it's much, much smaller than the multi-trillion-parameter models used in previous years. DeepSeek V4 Flash is the only model (so far) that can be run on small local GPU setups and still win an IMO gold. This is a huge testament to the capacity of open-weight models, both pushing the frontier of intelligence and offering the best price point.\n\nWe are calling it now that IMO 2027 problems will get solved by local LLMs running on your phone and the future of software and mathematics is [open-weight AI](https://cline.bot/cline-pass?ref=cline.ghost.io).", "url": "https://wpnews.pro/news/deepseek-wins-imo-gold-on-12-cents", "canonical_source": "https://cline.ghost.io/deepseek-wins-imo-gold-on-12-cents/", "published_at": "2026-08-26 20:44:00+00:00", "updated_at": "2026-08-26 20:48:23.947790+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["DeepSeek V4 Flash", "Cline", "Claude Fable 5", "GPT-5.6 Sol", "Kimi K3", "Qwen 3.6 35B A3B", "DeepSeek V4 Pro", "GLM 5.2"], "alternates": {"html": "https://wpnews.pro/news/deepseek-wins-imo-gold-on-12-cents", "markdown": "https://wpnews.pro/news/deepseek-wins-imo-gold-on-12-cents.md", "text": "https://wpnews.pro/news/deepseek-wins-imo-gold-on-12-cents.txt", "jsonld": "https://wpnews.pro/news/deepseek-wins-imo-gold-on-12-cents.jsonld"}}