Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study
A new arXiv paper (2608.18795v1) quantifies why majority voting among LLM samples can backfire on hard questions, introducing a pluralistic agreement index Gamma decomposed into a mechanical preferenc…