Skip to content
Review Open access

Label-Free Threshold Selection for Out-of-Distribution Detection in Liver CT Segmentation

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Results indicate that clinically meaningful failure detection can be derived from unlabeled data, and propose a label-free framework for calibrating OOD score thresholds.

Abstract

Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.

Read PDF

Similar papers

Preprint Sep 2026

ThreshGuide: Class-Aware Labeled-Guided Thresholding for Semi-Supervised 3D Abdominal Multi-Organ Segmentation

Pseudo-labeling is a strong paradigm for semi-supervised medical image segmentation, yet its effectiveness is highly sensitive to confidence thresholding. In abdominal multi-organ segmentation, a fixed global threshold is particularly suboptimal because organ classes differ substantially in size, appearance, and learni...

Hong-Yu Liu, Yin-Long Wang, Lu-Sha Li et al. · 0 citations
#small language model Preprint Aug 2026

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

This work proposes a label-free selection criterion built on SUDO, a framework for evaluating clinical AI systems without ground-truth annotations, and shows that AURCC can be used to rank a variety of vision-language models on chest X-ray classification across three inter-hospital shift scenarios, under zero-shot and...

J. Larrea, L. Mansilla, Enzo Ferrante · 0 citations
#artificial intelligence Preprint Aug 2026

Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift

Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. We calibrate per-organ thresholds for an AMOS-trained nnU-Net, audit transfer to RAOS, and estimate local re-certification cost using case-level voxel false-negative rate (FNR). The AMOS control passes, but $7/12$ organs exceed...

Souraj Adhikary, Negar Chabi, André Mastmeyer · 0 citations
Review Open access Aug 2026

Downstream-Aware Automated QC of Images and AI-Generated Segmentations.

A 3D deep learning model that classified MR volumes and their automated segmentation outputs into downstream Accept, Reject, and Rework categories could support clinical triage after automated segmentation to focus radiologist effort on flagged cases and expedited data curation in research studies where scan/segmentati...

Abigail E Green, Mrinal K Dhar, Adriana V. Gregory et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond In-Distribution Metrics: A Systematic Out-of-Distribution Evaluation of Congenital Heart Disease Segmentation

Congenital heart disease (CHD) diagnosis and surgical planning often require patient-specific 3D anatomical models, but manual segmentation is labor-intensive, particularly in complex anatomies. Although deep-learning methods can automate this process, they are typically evaluated in-distribution, despite clinically re...

Aniketh Vijesh, Shrisharanyan Vasu, Abhijit Ramesh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.