cd /news/large-language-models/contamination-inflates-scores-but-ra… · home topics large-language-models article
[ARTICLE · art-121079] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

A new study from arXiv (2609.02899v1) finds that benchmark contamination inflates absolute scores but rarely reorders large language model (LLM) leaderboards, with a rank correlation of 0.997 between standard and paraphrase-controlled rankings. Analyzing 47 public models and 74 finetuned models across ARC, GSM8K, HellaSwag, and MMLU, the authors show contamination is largely uniform, affecting only 3 of 188 model-by-benchmark cases, and recommend leaderboards report paraphrase-controlled rankings with confidence intervals.

read1 min views2 publishedSep 4, 2026

arXiv:2609.02899v1 Announce Type: new Abstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC, GSM8K, HellaSwag, MMLU), we first calibrate the measure against ground truth: it recovers injected contamination dose-responsively (a corrected effect of +0.187 accuracy points for test-set leakage) and never flags a negative-control model trained only on the legitimate training split (-0.012). We then quantify leaderboard impact: the rank correlation between a standard leaderboard and a paraphrase-controlled leaderboard is 0.997, and a sensitivity analysis shows that the observed differential contamination is far below the level needed to move rankings, with only 3 of 188 model-by-benchmark cases showing differential contamination corroborated across two references. Contamination among these public models is therefore largely uniform: it inflates absolute scores without reordering the leaderboard, and ranking distortion requires the rare case of differential contamination. We provide a calibrated invariance audit, released as a reference implementation, and recommend that leaderboards report paraphrase-controlled rankings alongside confidence intervals.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/contamination-inflat…] indexed:0 read:1min 2026-09-04 ·