cd /news/artificial-intelligence/deepseek-wins-imo-gold-on-12-cents · home topics artificial-intelligence article
[ARTICLE · art-112332] src=cline.ghost.io ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DeepSeek wins IMO Gold on 12 cents

DeepSeek V4 Flash scored 30/42 on IMO 2026 problems in Cline, clearing the 29-point gold medal cutoff for just $0.12, roughly 140 times cheaper than Claude Fable 5. The benchmark, run by Cline, tested eight models including GPT-5.6 Sol, which scored a perfect 42/42, and used blind grading by GPT-5.5 and Claude Opus 5 with Gemini 3.1 Pro resolving disagreements.

read6 min views1 publishedAug 26, 2026
DeepSeek wins IMO Gold on 12 cents
Image: Cline (auto-discovered)

Insights Benchmarking eight AI models on IMO 2026 problems in Cline, we found DeepSeek V4 Flash scored 30/42, clearing the gold medal cutoff for just $0.12. See how open-weight models compare with GPT-5.6 Sol, Claude Fable 5, Kimi K3, and more.

We wanted to know the cheapest possible way to win an IMO gold medal, so we ran eight models against IMO 2026 problems in Cline and had the proofs graded blind.

The answer was 12 cents.

Frontier models scoring perfect runs has been covered extensively . Our interesting find is that open-weight models now score gold too. Our DeepSeek V4 Flash run hit 30/42, clearing the 29-point gold cutoff. This was roughly 140 times cheaper than Claude Fable 5. In this post, we explore the problem setup and judging, and share all the traces from these runs.

Problem Setup #

The original benchmark was an 8 × 6 matrix: eight models and six IMO 2026 problems, making a total of 48 effective cells.

  • GPT-5.6 Sol
  • Claude Fable 5
  • Kimi K3
  • DeepSeek V4 Flash
  • Qwen 3.6 35B A3B
  • DeepSeek V4 Pro
  • GLM 5.2
  • MiMo V2.5 Pro

Each model received the problem statement and one instruction to submit its strongest complete final solution. A submission tool was used to finish the run.

After the original symmetric panel revealed some harness bugs, we fixed the tool-calling issues within the Cline harness to give every model its best shot at scoring, so that any lost points reflect actual model performance rather than tool-calling failures.

Judging

The proofs were graded anonymously on the IMO 0–7 scale. Two independent, model-blind graders, GPT‑5.5 and Claude Opus 5, scored each submission, and disagreements were resolved by Gemini 3.1 Pro. GPT-5.6 Sol remained the reference perfect scorer at 42/42. Other groups have independently reproduced that perfect score here, here, and here (although their harness settings differ from ours). The gold medal score cutoff this year was 29.

Each candidate proof was placed into an immutable anonymous batch. The graders saw the problem, proof, rubric, and comparison reference (including Lean solutions), but not the candidate model's identity, to avoid bias. They checked that the IMO scoring rules were followed and that the final results passed validation.

Reward-hacking guardrails included disabling internet use in the system prompt and exposing only the submit-solution tool. No prior context about the problems like Lean artifacts, rationales, etc. was provided, so each model had to rely on its own ability. We also checked traces to make sure no browser use was done. And since the IMO 2026 problems are new, they aren't in these models' training sets yet.

One caveat is that IMO problems, unlike a traditional coding benchmark, are not a stable population estimate: small changes like prompt tweaks, retries, and provider issues (for open-weight models) can shift the scores a bit. We ran multiple attempts for all models and picked the best score on each problem rather than using standard measures like Pass@k. Note that the prices in the table below reflect the single best run for each model, not the net cost of all the exploration needed to work through tool-calling issues.

In general, variance of results is always something to keep in mind with any benchmark. We used an LLM as a judge instead of unit tests (math olympiad problems often don’t come with unit tests) or a formal Lean grader. You can look at unofficial Lean solutions here.

As promised, we have attached all the solution traces from the different models here; these traces include prices, generation costs, and the solutions offered by each model.

Model Scores and Prices #

Model Best score P1–P6 Price
GPT-5.6 Sol 42/42
7 · 7 · 7 · 7 · 7 · 7 $3.2336
Claude Fable 5 41/42
7 · 7 · 6 · 7 · 7 · 7 $17.1956
Kimi K3 35/42
7 · 4 · 7 · 7 · 7 · 3 $5.1328
DeepSeek V4 Flash 30/42
7 · 4 · 3 · 7 · 7 · 2 $0.1215
DeepSeek V4 Pro 30/42
7 · 4 · 2 · 7 · 7 · 3 $0.4761
MiMo V2.5 Pro 30/42
7 · 2 · 3 · 7 · 7 · 4 $1.0650
GLM 5.2 21/42
7 · 0 · 0 · 7 · 7 · 0 $2.3855
Qwen 3.6 35B A3B 16/42
7 · 2 · 1 · 2 · 3 · 1 $0.2251
IMO 2026

16/42### History of AI models in IMO

Originally, in 2024, DeepMind's custom AlphaProof models (which were not released to the public) got an IMO silver, and back then the problems were converted to Lean. The models solved one problem within minutes and took up to three days to solve the others; note that humans solved the problems over two days, in sessions of 4.5 hours each.

In 2025, an advanced version of Gemini Deep Think achieved a gold medal score working directly from the official natural-language problem descriptions.

In 2026, we have made great progress in IMO and other math olympiad problem solving. For starters, we have migrated to much better mathematical benchmarks like Riemann bench, which are far more intricate and long-horizon, and models are now able to solve many unsolved Erdos problems. There are many other interesting solutions to math problems like Hadamard matrices and the Jacobian conjecture that were released this year, but there’s something very special that came from the open-weights community.

What's unique about DeepSeek V4 Flash

The most important part of the progress this year has been an order-of-magnitude improvement in the performance of open-weight models. Notably, DeepSeek V4 Flash isn’t post-trained specifically for Olympiad problems we hit the model endpoint directly and still got a gold medal. It has 284B total MoE parameters, with about 13B activated per token. While it's not exactly a small model, it's much, much smaller than the multi-trillion-parameter models used in previous years. DeepSeek V4 Flash is the only model (so far) that can be run on small local GPU setups and still win an IMO gold. This is a huge testament to the capacity of open-weight models, both pushing the frontier of intelligence and offering the best price point.

We are calling it now that IMO 2027 problems will get solved by local LLMs running on your phone and the future of software and mathematics is open-weight AI.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @deepseek v4 flash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-wins-imo-go…] indexed:0 read:6min 2026-08-26 ·