# We Checked 162 AI Benchmark Gaps. Only 20 Separate Cleanly. The Bigger Problem Was the Missing Data.

> Source: <https://dev.to/driftproofhq/we-checked-162-ai-benchmark-gaps-only-20-separate-cleanly-the-bigger-problem-was-the-missing-data-4d3e>
> Published: 2026-10-09 03:43:41+00:00

Six frontier AI launch posts and nine public leaderboards gave us 162 model-vs-model gaps to check.

20 separate cleanly.

That was not the result that bothered us most.

The bigger problem was how often the published data did not let us answer the question at all.

Of the 44 gaps quoted in the six launch posts, 11 separate, 5 do not, and 28 cannot be checked honestly from what was published.

Of 118 adjacent leaderboard pairs, 9 separate.

And across the broader 617-row source audit, 483 comparisons cannot be given an honest repeat-aware interval from the public information.

That is the finding.

Not “benchmark scores are fake.”

Not “these models are secretly equal.”

Something more basic:

AI benchmark charts often look more precise than the public evidence lets us verify.

Gemini 4 Argon’s launch material contains 15 relevant comparisons.

None supports an honest interval from the published data.

Mistral Large 4 contributes another 9.

Again, none.

The reasons vary: scores averaged across repeated runs without publishing the repeat count, different harnesses or effort settings on the two sides, or insufficient per-task results to reconstruct a paired comparison.

We did not manufacture precision to fill those gaps.

They are marked not checkable.

That distinction matters.

Not checkable does not mean wrong.

It means the evidence required to decide was not published.

SWE-bench Multilingual makes the problem easy to see.

The top two entries in our snapshot are:

72.7% vs 66.3%.

Looking at the leaderboard, a 6.4-point lead feels substantial.

But the benchmark contains 300 tasks.

Around the statistically least precise part of the scale, its conservative resolution is roughly 24 tasks. That is a benchmark-level reference point, not a universal cutoff: the exact interval depends on where the two scores sit.

For the actual 72.7% vs 66.3% pair, the applicable interval still includes zero.

The benchmark does not establish separation between them.

In fact, six positions on that leaderboard cannot be distinguished from first place under the applicable comparison.

That does not mean those models are equal.

It means something narrower:

this benchmark, at this size, with these published results, cannot separate them.

This is the part that made me trust the final result more.

Our frozen first pass produced 21 separated gaps.

Then we audited every comparison against the benchmark’s actual score type and any uncertainty information published by the source.

Five of our original “separated” verdicts did not survive that audit.

Four gaps that had looked unresolved under the simpler treatment did separate when the source’s own intervals were used.

Final count:

20.

The correction went both ways.

We did not overwrite the frozen results. The original verdict and the audited verdict sit next to each other in the dataset, with the reason for the change recorded.

There are two denominators in the data, so this is worth making explicit.

The 162 are the headline comparisons in this analysis:

The broader source audit contains 617 comparison rows, including an appendix of effort-curve contrasts that never entered the 162 headline count.

Of those 617 rows, 483 do not support an honest repeat-aware interval from the published information.

That result surprised us more than the 20/162 headline.

The replication problem often starts before statistical significance.

It starts with whether enough information was published to calculate uncertainty at all.

The analysis method was frozen and hashed before the benchmark scores were inspected.

For interpretable binary benchmark rates with a known number of questions, we use Wilson 95% intervals and a Newcombe hybrid Wilson interval for the difference.

Where valid per-task paired data exist, paired evidence takes precedence.

Where a source publishes suitable repeated-run uncertainty, that evidence takes precedence over pretending repeated runs are extra independent benchmark questions.

If the required information is missing, the answer is not guessed.

It is marked:

not calculable from published data.

The repository preserves the frozen method, source hashes, per-comparison results and the later score-type audit.

It does not prove that two models with overlapping intervals have equal ability.

It does not prove that every comparison without a separated verdict is false.

It does not prove that benchmark sampling noise captures every source of uncertainty in an AI evaluation.

And it definitely does not tell you which model will work better for your workload.

It answers one deliberately narrow question:

When someone puts two benchmark scores next to each other and calls one higher, does the published evidence establish that the gap is larger than the uncertainty we can actually calculate?

For these 162 comparisons, the answer was yes 20 times.

Full write-up and tables:

[https://driftproofhq.com/leaderboard-noise/](https://driftproofhq.com/leaderboard-noise/)

All 617 rows, frozen method, provenance and hashes:

[https://github.com/driftproofhq/ranknoise](https://github.com/driftproofhq/ranknoise)

And if you want to ignore our interpretation entirely and check a gap yourself:

[https://driftproofhq.com/benchmark-gap/](https://driftproofhq.com/benchmark-gap/)

Put in two scores and the number of benchmark questions. It returns the interval and expresses the difference in questions rather than only percentage points.

I maintain Driftproof, the open-source evaluation instrument this methodology grew out of.

Claude Code was used to build and run the analysis pipeline, including implementation of the interval calculations, source extraction, audit checks and reproducible result generation against the frozen method.

The data, calculator and methodology are public.

One question for people who actually use these benchmark tables to choose models:

What is the smallest benchmark lead you would personally accept as evidence that Model A is better than Model B without seeing repeated runs or per-task results?__
