cd /news/large-language-models/every-wrong-answer-counts-option-lev… · home topics large-language-models article
[ARTICLE · art-87118] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

Researchers introduced the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicted held-out interactions more accurately than binary models, achieving a Spearman correlation of 0.920 with the Arena.ai Elo leaderboard, and incorrect responses alone recovered ability estimates with Spearman 0.943. The framework also enabled efficient benchmarking, with 41 selected items preserving full-bank ranking with Kendall's correlation 0.85, a 770 times reduction.

read1 min views1 publishedAug 5, 2026

arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.

── more in #large-language-models 4 stories · sorted by recency
── more on @llm nominal response model 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/every-wrong-answer-c…] indexed:0 read:1min 2026-08-05 ·