{"slug": "when-agents-disagree-bayesian-backward-reasoning-as-a-label-free-anchor-for", "title": "When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making", "summary": "A new arXiv paper (2609.11709v1) proposes Bayesian backward reasoning as a label-free anchor for multi-agent collective decision-making, using Jensen-Shannon divergence to rank LLM agents by cross-path consistency. Evaluated on DDXPlus across five LLM backbones, the method's log-linear fusion strategy (LogLin) achieved the best performance among the evaluated methods, with its largest gains on the subset where agents disagree, while hard selection (MinJS) outperformed random selection across all backbones and soft reweighting (FwdJS) generally improved over the strongest baseline. The authors report that the reverse posterior, despite weaker standalone accuracy, serves as a more useful anchor than forward-only alternatives, and that a lightweight two-stage calibration can further refine it when labeled data are available.", "body_md": "arXiv:2609.11709v1 Announce Type: new \nAbstract: When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.", "url": "https://wpnews.pro/news/when-agents-disagree-bayesian-backward-reasoning-as-a-label-free-anchor-for", "canonical_source": "https://www.machinebrief.com/news/when-agents-disagree-bayesian-backward-reasoning-as-a-label-38y8", "published_at": "2026-09-12 04:00:00+00:00", "updated_at": "2026-09-12 05:27:47.424565+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "machine-learning"], "entities": ["arXiv", "DDXPlus", "MinJS", "FwdJS", "LogLin", "Jensen-Shannon divergence"], "alternates": {"html": "https://wpnews.pro/news/when-agents-disagree-bayesian-backward-reasoning-as-a-label-free-anchor-for", "markdown": "https://wpnews.pro/news/when-agents-disagree-bayesian-backward-reasoning-as-a-label-free-anchor-for.md", "text": "https://wpnews.pro/news/when-agents-disagree-bayesian-backward-reasoning-as-a-label-free-anchor-for.txt", "jsonld": "https://wpnews.pro/news/when-agents-disagree-bayesian-backward-reasoning-as-a-label-free-anchor-for.jsonld"}}