Skip to content
Review

Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

Jul 2026 · 0 citations · 28 references
Computer Science

Abstract

A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline. We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court found a Convention violation from the case's fact paragraphs. We compare three families across two frontier LLMs (Claude Opus 4.8 and GPT-5.5) as per-fact evidence estimators: (A) the raw LLM, (B) the LLM routed through the fusion pipeline, and (C) a term-frequency baseline through the same pipeline. Across roughly 4,750 tests we find: (1) on discrimination (AUROC around 0.83) the pipeline yields no improvement over either the raw LLM or the baseline; a frontier LLM used directly is the strongest single discriminator. (2) Naively composing an LLM with Bayesian-odds and Dempster-Shafer fusion more than doubles calibration error (ECE from about 0.16 to 0.46) via a prior-mismatch mechanism that replicates across both models. (3) Dempster-Shafer fusion is actively unsafe on long chains, committing confidently to wrong labels at below-chance accuracy; we recommend removing it. (4) The pipeline's genuine value is operational: routed through a conformal selective-prediction layer, the system decides which cases to automate and which to escalate. After removing Dempster-Shafer, recalibrating, and applying class-conditional risk control on the full 1,000-case set, the tuned engine auto-clears at 96.8 percent accuracy with 0.5 percent errors escaping and 96.3 percent caught for review, versus 85.9 / 3.8 / 72.1 for an untuned baseline. The contribution of such pipelines in law is calibrated trust, not sharper prediction.

View source

Similar papers

Preprint Aug 2026

Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a t...

Shu Yang · 0 citations
Preprint Apr 2026

When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation

The results suggest that directly asking for the target response may not always yield the most effective score for predicting it, and that comparing direct scores with indirect paths through related judgments may reveal a more effective predictive route.

Zonghuan Xu, Xiang Zheng, Yu-Tao Wu et al. · 0 citations
Preprint Aug 2026

Status Association Does Not Reliably Predict Decision Leakage

Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on...

X. Abdullah · 0 citations
Review Open access Aug 2026

The Physical Fidelity Gap as an Evidence-Traceability Problem in AI Uncertainty Quantification: A Structured Review

Artificial intelligence (AI) increasingly produces uncertainty outputs for sensing and measurement tasks, but the evidence supporting these outputs may not maintain a traceable correspondence with the relevant real-world conditions. This study conducted a structured 15-dimensional coding review of 566 studies, of which...

Lin Guo, Ai-Wen Ma, Heng Zhou et al. · 0 citations
Review Aug 2026

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.

Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting

This work proposes LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting), which reorganizes how collected evidence is used in the prediction stage and improves most prediction and calibration metrics across models and remains stronger under controlled comparisons of prior access, inference budget,...

Yu-Fei Chen, Yi-Ran Zhao, Xiao-Gang Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.