There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
A new arXiv study introduces the 'fragility grid,' showing that 12 open-weight instruction-tuned LLMs from 4 families score between 31 and 89 percent on the same 3,679 items from 4 benchmarks (ARC, He…