{"slug": "every-wrong-answer-counts-option-level-psychometrics-for-llm-multiple-choice", "title": "Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks", "summary": "Researchers introduced the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicted held-out interactions more accurately than binary models, achieving a Spearman correlation of 0.920 with the Arena.ai Elo leaderboard, and incorrect responses alone recovered ability estimates with Spearman 0.943. The framework also enabled efficient benchmarking, with 41 selected items preserving full-bank ranking with Kendall's correlation 0.85, a 770 times reduction.", "body_md": "arXiv:2608.02966v1 Announce Type: new\nAbstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.", "url": "https://wpnews.pro/news/every-wrong-answer-counts-option-level-psychometrics-for-llm-multiple-choice", "canonical_source": "https://arxiv.org/abs/2608.02966", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:04:13.719815+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning"], "entities": ["LLM Nominal Response Model", "Arena.ai"], "alternates": {"html": "https://wpnews.pro/news/every-wrong-answer-counts-option-level-psychometrics-for-llm-multiple-choice", "markdown": "https://wpnews.pro/news/every-wrong-answer-counts-option-level-psychometrics-for-llm-multiple-choice.md", "text": "https://wpnews.pro/news/every-wrong-answer-counts-option-level-psychometrics-for-llm-multiple-choice.txt", "jsonld": "https://wpnews.pro/news/every-wrong-answer-counts-option-level-psychometrics-for-llm-multiple-choice.jsonld"}}