cd /news/artificial-intelligence/calibrated-trust-not-sharper-predict… · home › topics › artificial-intelligence › article
[ARTICLE · art-100783] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

A study of 1,000 European Court of Human Rights cases found that fusing uncertainty tools into a legal AI pipeline does not improve prediction accuracy over a raw LLM, but can enhance operational trust. The pipeline, tested with Claude Opus 4.8 and GPT-5.5, doubled calibration error (ECE from 0.16 to 0.46) when naively composing LLM with Bayesian-odds and Dempster-Shafer fusion, and Dempster-Shafer was unsafe on long chains. After removing Dempster-Shafer and recalibrating, the tuned engine auto-cleared at 96.8% accuracy with 0.5% errors escaping and 96.3% caught for review, versus 85.9/3.8/72.1 for an untuned baseline.

read1 min views19 publishedAug 18, 2026

arXiv:2608.14617v1 Announce Type: new Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline. We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court found a Convention violation from the case's fact paragraphs. We compare three families across two frontier LLMs (Claude Opus 4.8 and GPT-5.5) as per-fact evidence estimators: (A) the raw LLM, (B) the LLM routed through the fusion pipeline, and (C) a term-frequency baseline through the same pipeline. Across roughly 4,750 tests we find: (1) on discrimination (AUROC around 0.83) the pipeline yields no improvement over either the raw LLM or the baseline; a frontier LLM used directly is the strongest single discriminator. (2) Naively composing an LLM with Bayesian-odds and Dempster-Shafer fusion more than doubles calibration error (ECE from about 0.16 to 0.46) via a prior-mismatch mechanism that replicates across both models. (3) Dempster-Shafer fusion is actively unsafe on long chains, committing confidently to wrong labels at below-chance accuracy; we recommend removing it. (4) The pipeline's genuine value is operational: routed through a conformal selective-prediction layer, the system decides which cases to automate and which to escalate. After removing Dempster-Shafer, recalibrating, and applying class-conditional risk control on the full 1,000-case set, the tuned engine auto-clears at 96.8 percent accuracy with 0.5 percent errors escaping and 96.3 percent caught for review, versus 85.9 / 3.8 / 72.1 for an untuned baseline. The contribution of such pipelines in law is calibrated trust, not sharper prediction.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @european court of human rights 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/calibrated-trust-not…] indexed:0 read:1min 2026-08-18 · —