cd /news/large-language-models/comed-the-missing-middle-between-rou… · home topics large-language-models article
[ARTICLE · art-138826] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference

Researchers introduced COMED (Controlled Model Escalation for Multi-LLM Deliberation), a post-anchor controller that selectively escalates queries to peer models instead of routing once or invoking peers on every query, according to an arXiv paper (2609.26913v1). COMED improved fixed and routed anchors in all 16 open-weight settings, with gains up to +10.7 percentage points on MedQA, and raised GPT-5.5 from 23.1% to 28.1% on HLE while invoking fewer models and decoding fewer tokens than dense collaboration. The paper formalizes the trade-off with a rescue-harm decomposition showing selective collaboration helps when rescued errors outweigh collaboration-induced harms.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.26913v1 Announce Type: new Abstract: No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model, while dense collaboration invokes peers on every query. We show that collaboration is non-monotonic: peers can recover failures that no model solves alone, but can also corrupt initially correct answers. We introduce COMED (Controlled Model Escalation for Multi-LLM Deliberation), a post-anchor controller for selective cross-model collaboration. COMED uses anchor self-consistency, router margin, and a lightweight peer probe to accept confident answers, verify ambiguous cases, and escalate only when collaboration is likely beneficial. We formalize this trade-off with a rescue-harm decomposition showing that selective collaboration improves when rescued errors outweigh collaboration-induced harms. Across medical, scientific, and general reasoning benchmarks, COMED improves fixed and routed anchors in all 16 open-weight settings, with gains up to +10.7 percentage points on MedQA while invoking fewer models and using fewer decoded tokens than dense collaboration. On HLE with frontier models, COMED improves GPT-5.5 from 23.1% to 28.1%, outperforming dense collaboration and achieving the best results.

── more in #large-language-models 4 stories · sorted by recency
── more on @comed 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/comed-the-missing-mi…] indexed:0 read:1min 2026-09-24 ·