{"slug": "we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was", "title": "We Checked 162 AI Benchmark Gaps. Only 20 Separate Cleanly. The Bigger Problem Was the Missing Data.", "summary": "An analysis of 162 model-vs-model gaps drawn from six frontier AI launch posts and nine public leaderboards found that only 20 separate cleanly, while 28 of the 44 gaps quoted in launch posts cannot be checked honestly from published data. Across a broader 617-row source audit, 483 comparisons lack enough published information to compute an honest repeat-aware interval, a problem the authors say often precedes any question of statistical significance. The team froze and hashed its method before inspecting scores, and five of its original \"separated\" verdicts failed a later audit while four unresolved gaps separated once sources' own intervals were used.", "body_md": "Six frontier AI launch posts and nine public leaderboards gave us 162 model-vs-model gaps to check.\n\n20 separate cleanly.\n\nThat was not the result that bothered us most.\n\nThe bigger problem was how often the published data did not let us answer the question at all.\n\nOf the 44 gaps quoted in the six launch posts, 11 separate, 5 do not, and 28 cannot be checked honestly from what was published.\n\nOf 118 adjacent leaderboard pairs, 9 separate.\n\nAnd across the broader 617-row source audit, 483 comparisons cannot be given an honest repeat-aware interval from the public information.\n\nThat is the finding.\n\nNot “benchmark scores are fake.”\n\nNot “these models are secretly equal.”\n\nSomething more basic:\n\nAI benchmark charts often look more precise than the public evidence lets us verify.\n\nGemini 4 Argon’s launch material contains 15 relevant comparisons.\n\nNone supports an honest interval from the published data.\n\nMistral Large 4 contributes another 9.\n\nAgain, none.\n\nThe reasons vary: scores averaged across repeated runs without publishing the repeat count, different harnesses or effort settings on the two sides, or insufficient per-task results to reconstruct a paired comparison.\n\nWe did not manufacture precision to fill those gaps.\n\nThey are marked not checkable.\n\nThat distinction matters.\n\nNot checkable does not mean wrong.\n\nIt means the evidence required to decide was not published.\n\nSWE-bench Multilingual makes the problem easy to see.\n\nThe top two entries in our snapshot are:\n\n72.7% vs 66.3%.\n\nLooking at the leaderboard, a 6.4-point lead feels substantial.\n\nBut the benchmark contains 300 tasks.\n\nAround the statistically least precise part of the scale, its conservative resolution is roughly 24 tasks. That is a benchmark-level reference point, not a universal cutoff: the exact interval depends on where the two scores sit.\n\nFor the actual 72.7% vs 66.3% pair, the applicable interval still includes zero.\n\nThe benchmark does not establish separation between them.\n\nIn fact, six positions on that leaderboard cannot be distinguished from first place under the applicable comparison.\n\nThat does not mean those models are equal.\n\nIt means something narrower:\n\nthis benchmark, at this size, with these published results, cannot separate them.\n\nThis is the part that made me trust the final result more.\n\nOur frozen first pass produced 21 separated gaps.\n\nThen we audited every comparison against the benchmark’s actual score type and any uncertainty information published by the source.\n\nFive of our original “separated” verdicts did not survive that audit.\n\nFour gaps that had looked unresolved under the simpler treatment did separate when the source’s own intervals were used.\n\nFinal count:\n\n20.\n\nThe correction went both ways.\n\nWe did not overwrite the frozen results. The original verdict and the audited verdict sit next to each other in the dataset, with the reason for the change recorded.\n\nThere are two denominators in the data, so this is worth making explicit.\n\nThe 162 are the headline comparisons in this analysis:\n\nThe broader source audit contains 617 comparison rows, including an appendix of effort-curve contrasts that never entered the 162 headline count.\n\nOf those 617 rows, 483 do not support an honest repeat-aware interval from the published information.\n\nThat result surprised us more than the 20/162 headline.\n\nThe replication problem often starts before statistical significance.\n\nIt starts with whether enough information was published to calculate uncertainty at all.\n\nThe analysis method was frozen and hashed before the benchmark scores were inspected.\n\nFor interpretable binary benchmark rates with a known number of questions, we use Wilson 95% intervals and a Newcombe hybrid Wilson interval for the difference.\n\nWhere valid per-task paired data exist, paired evidence takes precedence.\n\nWhere a source publishes suitable repeated-run uncertainty, that evidence takes precedence over pretending repeated runs are extra independent benchmark questions.\n\nIf the required information is missing, the answer is not guessed.\n\nIt is marked:\n\nnot calculable from published data.\n\nThe repository preserves the frozen method, source hashes, per-comparison results and the later score-type audit.\n\nIt does not prove that two models with overlapping intervals have equal ability.\n\nIt does not prove that every comparison without a separated verdict is false.\n\nIt does not prove that benchmark sampling noise captures every source of uncertainty in an AI evaluation.\n\nAnd it definitely does not tell you which model will work better for your workload.\n\nIt answers one deliberately narrow question:\n\nWhen someone puts two benchmark scores next to each other and calls one higher, does the published evidence establish that the gap is larger than the uncertainty we can actually calculate?\n\nFor these 162 comparisons, the answer was yes 20 times.\n\nFull write-up and tables:\n\n[https://driftproofhq.com/leaderboard-noise/](https://driftproofhq.com/leaderboard-noise/)\n\nAll 617 rows, frozen method, provenance and hashes:\n\n[https://github.com/driftproofhq/ranknoise](https://github.com/driftproofhq/ranknoise)\n\nAnd if you want to ignore our interpretation entirely and check a gap yourself:\n\n[https://driftproofhq.com/benchmark-gap/](https://driftproofhq.com/benchmark-gap/)\n\nPut in two scores and the number of benchmark questions. It returns the interval and expresses the difference in questions rather than only percentage points.\n\nI maintain Driftproof, the open-source evaluation instrument this methodology grew out of.\n\nClaude Code was used to build and run the analysis pipeline, including implementation of the interval calculations, source extraction, audit checks and reproducible result generation against the frozen method.\n\nThe data, calculator and methodology are public.\n\nOne question for people who actually use these benchmark tables to choose models:\n\nWhat is the smallest benchmark lead you would personally accept as evidence that Model A is better than Model B without seeing repeated runs or per-task results?__", "url": "https://wpnews.pro/news/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was", "canonical_source": "https://dev.to/driftproofhq/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was-the-missing-data-4d3e", "published_at": "2026-10-09 03:43:41+00:00", "updated_at": "2026-10-09 03:47:56.833080+00:00", "lang": "en", "topics": ["ai-research", "machine-learning", "large-language-models", "artificial-intelligence"], "entities": ["Gemini 4 Argon", "Mistral Large 4", "SWE-bench Multilingual"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was", "markdown": "https://wpnews.pro/news/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was.md", "text": "https://wpnews.pro/news/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was.txt", "jsonld": "https://wpnews.pro/news/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was.jsonld"}}