cd /news/large-language-models/there-is-no-neutral-harness-modern-l… · home topics large-language-models article
[ARTICLE · art-109622] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

A new arXiv study introduces the 'fragility grid,' showing that 12 open-weight instruction-tuned LLMs from 4 families score between 31 and 89 percent on the same 3,679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) depending on 26 equally defensible harness configurations. Config-fragile items carry 95.7 percent of the score gap between adjacent models on average, and 4 of the 12 models reach rank one under some configuration, meaning the harness selects the winner. The authors release per-item records and analysis scripts for reproducibility.

read2 min views2 publishedAug 25, 2026

arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items. We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs from 4 families answer the same 3{,}679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) under 26 equally defensible harness configurations, recording one correctness bit for every model, item, and configuration. The comparison is matched, since the items, the weights, and the greedy decoding stay fixed while only the harness varies. Under the grid a model's score is a band rather than a point: gemma4-31b scores between 31 and 89 percent depending only on the harness. Three results follow. On the items that two adjacent models both answer stably the pair is tied, and config-fragile items carry 95.7 percent of a pair's gap on average. Four of the 12 models reach rank one under some configuration, so the harness selects the winner. Item discrimination, the property that benchmark-compression methods maximize, correlates with fragility at 0.28 (95 percent CI 0.25 to 0.30), so compression keeps the fragile items rather than removing them. The scoring choice, not the option order that protocols usually fix, is the load-bearing axis. We release the per-item records and the analysis script, from which every number regenerates on a CPU in seconds, and we position the fragility grid as a check a leaderboard can run before it reports an order.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/there-is-no-neutral-…] indexed:0 read:2min 2026-08-25 ·