Skip to content
Open access

CBEC: a simple retrieval-based framework for population-grounded prediction reliability estimation in chest X-ray diagnosis

Aug 2026 · Scientific Reports · 0 citations

TL;DR

It is suggested that case-level consistency provides a meaningful practical signal for reliability estimation in medical imaging, whose relationship to the classical epistemic–aleatoric decomposition warrants further theoretical investigation across broader clinical settings.

Abstract

Automated chest X-ray diagnosis fails most critically when models produce confident yet incorrect predictions, suppressing clinical oversight at the point of decision-making. Existing uncertainty estimation methods define epistemic uncertainty primarily as a property of model parameters, overlooking whether predictions remain consistent with clinically similar cases. We propose that the divergence between a model’s prediction and the empirical label distribution of neighbouring training cases provides a practical reliability signal—one that correlates with, and can serve as a proxy for, epistemic uncertainty, while acknowledging that it may also reflect additional sources of discrepancy including label noise, representation error, and local population variability. To operationalize this perspective, we introduce a retrieval-based framework that constructs a fixed embedding-space memory of training cases and estimates a non-parametric label distribution over nearest neighbours. This enables direct comparison between model predictions and local case-level structure without introducing additional trainable parameters or modifying the diagnostic backbone. Experiments on ChestX-ray14 demonstrate improved detection of misclassified predictions relative to deep ensembles, with approximately 2.9 percentage-point gains in uncertainty-based error identification and an approximately 23% reduction in calibration error. Under zero-shot transfer to PadChest and CheXpert, the proposed approach exhibits smaller degradation in uncertainty estimation quality, with larger improvements observed for rare pathological conditions. These findings suggest that case-level consistency provides a meaningful practical signal for reliability estimation in medical imaging, whose relationship to the classical epistemic–aleatoric decomposition warrants further theoretical investigation across broader clinical settings.

Read PDF

Similar papers

#machine learning Review Sep 2026

Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models

Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion tha...

Mingcheng Zhu, Jin-Ning Liang, Ting-Ting Zhu · 0 citations
Open access Sep 2026

Accuracy Overstates Evidence Grounding and Abstention Reliability in Mammography Vision-Language Models

An evidence-grounded selective evaluation benchmark that evaluates Pathology classification and Abnormality identification together with label-aware lesion localization suggests that answer accuracy and abstention behavior can substantially overstate the reliability of current VLM when predictions are not verified agai...

B. Qu, W. Liu, M. Murrow et al. · 0 citations
#natural language process... Preprint Sep 2026

How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification

A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence calibration metrics target binary or multi-class tasks, while multi-label calibration remains largely underexplored. Mult...

Sophie Henning, Georg Hofmann, Alexander Schulte et al. · 0 citations
Preprint Aug 2026

ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imaging data and task-specific training. We investigate whether a generalist large language model (LLM), reading only a faithful natural-language rendering of standard nodule...

Md. Rabiul Islam, Samir Abdaljalil, E. Serpedin et al. · 0 citations
#machine learning Preprint Aug 2026

Uncertainty of Vision Medical Foundation Models

This work underscores the need for a holistic approach to uncertainty quantification in recent development of medical vision foundation model, ensuring robust and interpretable AI-driven decision-making and highlights the importance of careful model selection and the inte- gration of both point and region prediction to...

Hao-Xu Huang, Narges Razavian · 0 citations
Preprint Aug 2026

MC-CXR: A Multi-Context Chest X-ray Benchmark for Context-Induced Disruption in Vision-Language Models

Vision-language models (VLMs) are increasingly used in clinical pipelines where a chest X-ray is interpreted alongside retrieved reports, preliminary notes, or prior imaging. Existing benchmarks measure whether models answer correctly in isolation, but not whether they preserve a correct image-only decision when plausi...

Junhyeok Lee, Songsoo Kim, Kyu Sung Choi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.