cd /news/large-language-models/uncertainty-aware-trust-estimation-f… · home topics large-language-models article
[ARTICLE · art-71397] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

A new study from arXiv introduces an uncertainty-aware trust estimation method for multi-LLM systems, adapting Cooke-style log weighting from structured expert judgment to aggregate predictions from heterogeneous LLMs. The approach, evaluated on MMLU and MMLU-Pro, achieves superior accuracy-reliability balance and robustness under contamination, outperforming naive aggregation methods that assume equal trustworthiness.

read1 min views1 publishedJul 24, 2026

arXiv:2607.20529v1 Announce Type: new Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality. This assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable or adversarial experts. In this work, we formulate multi-LLM aggregation as a problem of uncertainty-aware trust estimation. We adapt structured expert judgment from decision theory, using context-aware calibration questions to estimate expert reliability based on the quality of its probabilistic predictions. Specifically, we employ Cooke-style log weighting, which penalises overconfident incorrect predictions and favours well-calibrated experts. We evaluate our approach on MMLU and MMLU-Pro across homogeneous, heterogeneous, and contaminated expert panels. Results show that while aggregation methods perform similarly in homogeneous settings, Cooke weighting becomes critical under heterogeneity and contamination. It achieves a superior accuracy-reliability balance and remains robust when unreliable experts are introduced. These findings suggest that Multi-LLM aggregation requires not just combining predictions, but calibrating trust under uncertainty.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/uncertainty-aware-tr…] indexed:0 read:1min 2026-07-24 ·