It is suggested that case-level consistency provides a meaningful practical signal for reliability estimation in medical imaging, whose relationship to the classical epistemic–aleatoric decomposition warrants further theoretical investigation across broader clinical settings.
Abstract
Automated chest X-ray diagnosis fails most critically when models produce confident yet incorrect predictions, suppressing clinical oversight at the point of decision-making. Existing uncertainty estimation methods define epistemic uncertainty primarily as a property of model parameters, overlooking whether predictions remain consistent with clinically similar cases. We propose that the divergence between a model’s prediction and the empirical label distribution of neighbouring training cases provides a practical reliability signal—one that correlates with, and can serve as a proxy for, epistemic uncertainty, while acknowledging that it may also reflect additional sources of discrepancy including label noise, representation error, and local population variability. To operationalize this perspective, we introduce a retrieval-based framework that constructs a fixed embedding-space memory of training cases and estimates a non-parametric label distribution over nearest neighbours. This enables direct comparison between model predictions and local case-level structure without introducing additional trainable parameters or modifying the diagnostic backbone. Experiments on ChestX-ray14 demonstrate improved detection of misclassified predictions relative to deep ensembles, with approximately 2.9 percentage-point gains in uncertainty-based error identification and an approximately 23% reduction in calibration error. Under zero-shot transfer to PadChest and CheXpert, the proposed approach exhibits smaller degradation in uncertainty estimation quality, with larger improvements observed for rare pathological conditions. These findings suggest that case-level consistency provides a meaningful practical signal for reliability estimation in medical imaging, whose relationship to the classical epistemic–aleatoric decomposition warrants further theoretical investigation across broader clinical settings.
Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion tha...
An evidence-grounded selective evaluation benchmark that evaluates Pathology classification and Abnormality identification together with label-aware lesion localization suggests that answer accuracy and abstention behavior can substantially overstate the reliability of current VLM when predictions are not verified agai...
B. Qu, W. Liu, M. Murrow et al.· medRxiv· 0 citations
A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence calibration metrics target binary or multi-class tasks, while multi-label calibration remains largely underexplored. Mult...
Sophie Henning, Georg Hofmann, Alexander Schulte et al.· 0 citations
Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imaging data and task-specific training. We investigate whether a generalist large language model (LLM), reading only a faithful natural-language rendering of standard nodule...
Md. Rabiul Islam, Samir Abdaljalil, E. Serpedin et al.· 0 citations
This work underscores the need for a holistic approach to uncertainty quantification in recent development of medical vision foundation model, ensuring robust and interpretable AI-driven decision-making and highlights the importance of careful model selection and the inte- gration of both point and region prediction to...
Vision-language models (VLMs) are increasingly used in clinical pipelines where a chest X-ray is interpreted alongside retrieved reports, preliminary notes, or prior imaging. Existing benchmarks measure whether models answer correctly in isolation, but not whether they preserve a correct image-only decision when plausi...
Junhyeok Lee, Songsoo Kim, Kyu Sung Choi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.