Skip to content
Review

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

Claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.

Abstract

Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness. Response-level confidence is a coarse signal: a single generation can mix correct and incorrect statements, so a single number is not actionable for users that must accept, reject, or verify individual pieces of information. We study claim-level confidence calibration as a decision-relevant uncertainty signal: each response is decomposed into atomic, verifiable claims, and each claim is assigned a calibrated confidence using inference-time signals from consistency across samples and self-verification. Our framework operates in closed-box settings (no logits, no fine-tuning) and applies post-hoc calibration directly at the claim level, enabling selective intervention such as evidence retrieval or human review for low-confidence claims. Across TriviaQA and TruthfulQA we evaluate seven baselines on six recent models (Llama-3.1, Mistral, Qwen2.5, DeepSeek-R1, GPT-4, GPT-4o), and show that claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.

View source

Similar papers

Review Jul 2026

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

ConfidenceBench, a calibration benchmark that evaluates verbalized confidence estimates in 15 frontier LLMs using the Brier score, a proper scoring rule that incentivises truthful probability reporting shows that verbalized confidence calibration is a distinct and practically important axis of LLM reliability, complementary to standard accuracy-based evaluation.

M. ffrench-Constant, Daniel Yang, Xinmeng Huang et al. · 1 citation · ⚡1
Jul 2026

Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees

Experimental results show that CIC consistently achieves valid risk control while retaining strong answering efficiency, providing a practical and statistically grounded mechanism for deploying LLMs in reliability-sensitive QA workflows.

Sijin Dong, Hiroyuki Shinnou · 0 citations
Jul 2026

The Computational Basis of Confidence in Large Language Models

A computational account of confidence in multimodal language models is provided, when answer logits behave as readouts of a latent decision variable is delineated, and statistical decision confidence is established as a unifying framework for studying confidence across biological and artificial intelligence.

D. Kumaran, Viorica Patraucean, M. Ovsjanikov et al. · 0 citations
Preprint Aug 2026

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

DirEAG is proposed, a Dirichlet Evidence Aggregation method that converts each elicited answer-confidence observation into calibrated soft evidence over generated candidate answers and an additional null state, allowing the model to represent cases where none of the candidates is correct.

Haorui Xu, Yu-Zhou Zhu, Liyuan Gao · 0 citations
#artificial intelligence Preprint Aug 2026

When Linguistic and Internal Confidence Diverge in Large Language Models

Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls, and support a lossy-channel view of linguistic confidence.

Hefan Zhang, Bing-Quan Zhang, Ming Cheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.