{"slug": "contamination-inflates-scores-but-rarely-reorders-large-language-model", "title": "Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards", "summary": "A new study from arXiv (2609.02899v1) finds that benchmark contamination inflates absolute scores but rarely reorders large language model (LLM) leaderboards, with a rank correlation of 0.997 between standard and paraphrase-controlled rankings. Analyzing 47 public models and 74 finetuned models across ARC, GSM8K, HellaSwag, and MMLU, the authors show contamination is largely uniform, affecting only 3 of 188 model-by-benchmark cases, and recommend leaderboards report paraphrase-controlled rankings with confidence intervals.", "body_md": "arXiv:2609.02899v1 Announce Type: new\nAbstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC, GSM8K, HellaSwag, MMLU), we first calibrate the measure against ground truth: it recovers injected contamination dose-responsively (a corrected effect of +0.187 accuracy points for test-set leakage) and never flags a negative-control model trained only on the legitimate training split (-0.012). We then quantify leaderboard impact: the rank correlation between a standard leaderboard and a paraphrase-controlled leaderboard is 0.997, and a sensitivity analysis shows that the observed differential contamination is far below the level needed to move rankings, with only 3 of 188 model-by-benchmark cases showing differential contamination corroborated across two references. Contamination among these public models is therefore largely uniform: it inflates absolute scores without reordering the leaderboard, and ranking distortion requires the rare case of differential contamination. We provide a calibrated invariance audit, released as a reference implementation, and recommend that leaderboards report paraphrase-controlled rankings alongside confidence intervals.", "url": "https://wpnews.pro/news/contamination-inflates-scores-but-rarely-reorders-large-language-model", "canonical_source": "https://arxiv.org/abs/2609.02899", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:22:14.518863+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv", "ARC", "GSM8K", "HellaSwag", "MMLU"], "alternates": {"html": "https://wpnews.pro/news/contamination-inflates-scores-but-rarely-reorders-large-language-model", "markdown": "https://wpnews.pro/news/contamination-inflates-scores-but-rarely-reorders-large-language-model.md", "text": "https://wpnews.pro/news/contamination-inflates-scores-but-rarely-reorders-large-language-model.txt", "jsonld": "https://wpnews.pro/news/contamination-inflates-scores-but-rarely-reorders-large-language-model.jsonld"}}