{"slug": "decomposing-wrong-consensus-agreement-in-llm-self-consistency-a-gpt-4-1-case", "title": "Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study", "summary": "A new arXiv paper (2608.18795v1) quantifies why majority voting among LLM samples can backfire on hard questions, introducing a pluralistic agreement index Gamma decomposed into a mechanical preference component and a residual. On GPT-4.1, the per-case answer preference explains 81-93% of the agreement index on GPQA-Diamond, but only 59-78% on AIME, leaving a residual of 1.56-2.80 Gamma units. The study reproduces a voting gap down to -0.09 on hard questions and finds that the highest-agreement bin reaches only 0.42-0.83 accuracy, concluding agreement is graded evidence, not certification.", "body_md": "arXiv:2608.18795v1 Announce Type: new\nAbstract: Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative account of this failure. A pluralistic agreement index Gamma is defined as the expected fraction of the samples of a wrong run that agree with the consensus, normalized by a reference scale d=(1-p)/(C-1), and is decomposed into a mechanical component (what a vote delivers given only a per-case answer preference) and a preference-unexplained residual. The mechanical null is difficulty-matched and leak-free: each case is resimulated at its own accuracy and option preference, estimated from the case's other runs, so no run predicts its own agreement. On GPT-4.1 the decomposition shows benchmark-associated direction (an observational ordering over n=4 cells per benchmark, not a significance claim). On multiple-choice GPQA-Diamond, the per-case answer preference explains 81-93% of the held-out test-run agreement index: the shared-bias-dominates account over-claims here, because a wrong but attractive option the whole cohort latches onto is captured by the per-case preference channel (whether that preference is induced by shared training bias is not identified). On open-domain AIME, the mechanical preference explains only 59-78% (21-29% if shrunk to pure noise), and a preference-unexplained residual of 1.56-2.80 Gamma units survives, which a run-level preference-heterogeneity reference more than absorbs (1.4-2.1). A self-consistency backfire on hard questions is reproduced (binned voting gap down to -0.09, coupled CI [-0.12,-0.07]), and the highest-agreement bin reaches an accuracy of only 0.42-0.83, a 1.2-3.6x lift over base rate: agreement is graded evidence, not certification. No new voting method is proposed; code and evidence are committed and reproducible.", "url": "https://wpnews.pro/news/decomposing-wrong-consensus-agreement-in-llm-self-consistency-a-gpt-4-1-case", "canonical_source": "https://www.machinebrief.com/news/decomposing-wrong-consensus-agreement-in-llm-self-consistenc-tzvz", "published_at": "2026-08-20 04:00:00+00:00", "updated_at": "2026-08-20 05:15:01.348745+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence"], "entities": ["GPT-4.1", "GPQA-Diamond", "AIME", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/decomposing-wrong-consensus-agreement-in-llm-self-consistency-a-gpt-4-1-case", "markdown": "https://wpnews.pro/news/decomposing-wrong-consensus-agreement-in-llm-self-consistency-a-gpt-4-1-case.md", "text": "https://wpnews.pro/news/decomposing-wrong-consensus-agreement-in-llm-self-consistency-a-gpt-4-1-case.txt", "jsonld": "https://wpnews.pro/news/decomposing-wrong-consensus-agreement-in-llm-self-consistency-a-gpt-4-1-case.jsonld"}}