cd /news/artificial-intelligence/reasoning-consensus-structural-ensem… · home topics artificial-intelligence article
[ARTICLE · art-81345] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

Researchers propose a framework that ensembles the reasoning structure of multiple large language models (LLMs) by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains, weighting each step by how many traces independently attest to it to return 'Consensus Reasoning'. Across six benchmarks, the ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM, and consensus subgraphs are preferred over alternatives in 54.4-65.4% of head-to-head comparisons across five of six datasets.

read1 min views1 publishedJul 31, 2026

arXiv:2607.27783v1 Announce Type: new Abstract: Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return "Consensus Reasoning". Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman $\rho = 0.30$-$0.51$, and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4-65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reasoning-consensus-…] indexed:0 read:1min 2026-07-31 ·