{"slug": "calibrated-trust-not-sharper-prediction-an-empirical-test-of-uncertainty-fusion", "title": "Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion", "summary": "A study of 1,000 European Court of Human Rights cases found that fusing uncertainty tools into a legal AI pipeline does not improve prediction accuracy over a raw LLM, but can enhance operational trust. The pipeline, tested with Claude Opus 4.8 and GPT-5.5, doubled calibration error (ECE from 0.16 to 0.46) when naively composing LLM with Bayesian-odds and Dempster-Shafer fusion, and Dempster-Shafer was unsafe on long chains. After removing Dempster-Shafer and recalibrating, the tuned engine auto-cleared at 96.8% accuracy with 0.5% errors escaping and 96.3% caught for review, versus 85.9/3.8/72.1 for an untuned baseline.", "body_md": "arXiv:2608.14617v1 Announce Type: new\nAbstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline. We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court found a Convention violation from the case's fact paragraphs. We compare three families across two frontier LLMs (Claude Opus 4.8 and GPT-5.5) as per-fact evidence estimators: (A) the raw LLM, (B) the LLM routed through the fusion pipeline, and (C) a term-frequency baseline through the same pipeline. Across roughly 4,750 tests we find: (1) on discrimination (AUROC around 0.83) the pipeline yields no improvement over either the raw LLM or the baseline; a frontier LLM used directly is the strongest single discriminator. (2) Naively composing an LLM with Bayesian-odds and Dempster-Shafer fusion more than doubles calibration error (ECE from about 0.16 to 0.46) via a prior-mismatch mechanism that replicates across both models. (3) Dempster-Shafer fusion is actively unsafe on long chains, committing confidently to wrong labels at below-chance accuracy; we recommend removing it. (4) The pipeline's genuine value is operational: routed through a conformal selective-prediction layer, the system decides which cases to automate and which to escalate. After removing Dempster-Shafer, recalibrating, and applying class-conditional risk control on the full 1,000-case set, the tuned engine auto-clears at 96.8 percent accuracy with 0.5 percent errors escaping and 96.3 percent caught for review, versus 85.9 / 3.8 / 72.1 for an untuned baseline. The contribution of such pipelines in law is calibrated trust, not sharper prediction.", "url": "https://wpnews.pro/news/calibrated-trust-not-sharper-prediction-an-empirical-test-of-uncertainty-fusion", "canonical_source": "https://arxiv.org/abs/2608.14617", "published_at": "2026-08-18 04:00:00+00:00", "updated_at": "2026-08-18 04:12:31.806170+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-safety"], "entities": ["European Court of Human Rights", "LexGLUE", "FairLex", "Claude Opus 4.8", "GPT-5.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/calibrated-trust-not-sharper-prediction-an-empirical-test-of-uncertainty-fusion", "markdown": "https://wpnews.pro/news/calibrated-trust-not-sharper-prediction-an-empirical-test-of-uncertainty-fusion.md", "text": "https://wpnews.pro/news/calibrated-trust-not-sharper-prediction-an-empirical-test-of-uncertainty-fusion.txt", "jsonld": "https://wpnews.pro/news/calibrated-trust-not-sharper-prediction-an-empirical-test-of-uncertainty-fusion.jsonld"}}