{"slug": "llms-can-predict-failure-risk-but-struggle-to-predict-which-collaboration-pays", "title": "LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks", "summary": "A new arXiv study (2608.14927v1) finds that large language models can predict failure risk but struggle to identify which multi-agent collaboration protocol is cost-effective. Across 4,181 competition-level math problems, a post-answer gpt-oss-120b probe achieved 0.8847 AUROC for predicting Baseline failures, but only 0.1674 and 0.1041 AUPRC for identifying PER- or Broadcast-specific value. The pre-answer self-confidence gate reached 78.0% solve at 45K tokens, versus 73.8% at 71.3K for a frozen gpt-oss-120b router, while a retrospective oracle added 23.2-58.3 points of coverage over Baseline across 10 settings.", "body_md": "arXiv:2608.14927v1 Announce Type: new\nAbstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while holding the solver fixed within each setting: direct solving (Baseline), iterative self-correction (Single), planner-executor-reviewer collaboration (PER), and multi-agent deliberation (Broadcast). The primary benchmark comprises 4,181 competition-level math problems; paired robustness checks cover four benchmarks spanning competition math, biology, and broader science with two solver families. Across fixed policies, trained routers, and frozen LLM routers, conservative policies under-escalate, whereas higher-solve frozen routers often over-escalate. A post-answer, pre-collaboration gpt-oss-120b probe ranks Baseline failures with 0.8847 AUROC (4,151 parseable cases; 95% CI [0.8732, 0.8955]). The same score remains informative for predicting whether any collaboration helps (0.7683 AUPRC), but is much weaker for identifying PER- or Broadcast-specific value (0.1674 and 0.1041 AUPRC). Separately, the pre-answer self-confidence gate reaches 78.0% solve at 45K tokens, compared with 73.8% at 71.3K for a frozen gpt-oss-120b router and 92.4% for a retrospective fixed-order oracle. Across 10 paired model-condition settings, the oracle adds 23.2-58.3 points of retrospective coverage over Baseline, but protocol profiles vary by task. In the six settings with held-out router evaluations, oracle gaps remain 18.5-28.9 points. Confidence can therefore support initial escalation, while protocol-specific cost-aware routing remains unresolved.", "url": "https://wpnews.pro/news/llms-can-predict-failure-risk-but-struggle-to-predict-which-collaboration-pays", "canonical_source": "https://www.machinebrief.com/news/llms-can-predict-failure-risk-but-struggle-to-predict-which-y0j1", "published_at": "2026-08-18 04:00:00+00:00", "updated_at": "2026-08-18 04:41:03.218785+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents"], "entities": ["arXiv", "gpt-oss-120b"], "alternates": {"html": "https://wpnews.pro/news/llms-can-predict-failure-risk-but-struggle-to-predict-which-collaboration-pays", "markdown": "https://wpnews.pro/news/llms-can-predict-failure-risk-but-struggle-to-predict-which-collaboration-pays.md", "text": "https://wpnews.pro/news/llms-can-predict-failure-risk-but-struggle-to-predict-which-collaboration-pays.txt", "jsonld": "https://wpnews.pro/news/llms-can-predict-failure-risk-but-struggle-to-predict-which-collaboration-pays.jsonld"}}