Skip to content
Preprint

How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

Aug 2026 · 1 citation · 52 references
Computer Science

TL;DR

This study finds that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable.

Abstract

Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, with ten UDA algorithms and 13 label-free selection methods (validators), evaluating over 80,000 trained models in total. By this, we find that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable. Towards closing it, we explore two strategies, ensembling and a small target-labeling budget; both narrow this gap but do not close it entirely. Overall, deployable UDA depends on the complete pipeline; addressing the less explored selection step could bring much of current UDA closer to clinical use.

View source

Similar papers

Preprint Aug 2026

How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification

Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, practical deployment requires committing to a sampling strategy before the full annotation budget is spent, and choosing the wrong strategy can increase rather than decrease costs. W...

Julia Machnio, Mads Nielsen, M. M. Ghazi · 0 citations
#small language model Preprint Aug 2026

Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy

This work proposes a label-free selection criterion built on SUDO, a framework for evaluating clinical AI systems without ground-truth annotations, and shows that AURCC can be used to rank a variety of vision-language models on chest X-ray classification across three inter-hospital shift scenarios, under zero-shot and...

J. Larrea, L. Mansilla, Enzo Ferrante · 0 citations
Open access Aug 2026

Cross-Domain Generalization Neural Architecture Search for Robust Clinical Image Analysis

A Cross-Domain Generalization Neural Architecture Search (CDG-NAS) framework that automatically discovers neural network architectures capable of maintaining high diagnostic accuracy under diverse clinical domain shifts and incorporates domain robustness as a first-class optimization objective throughout the search pro...

A. Devi, L. Atlas · 0 citations
Preprint Aug 2026

Conformal Risk Minimization for Semi-Supervised Domain Adaptation via Optimal Transport

In high-stakes healthcare applications, machine learning models are frequently trained on data from one patient population and deployed on another, creating a distribution shift that degrades both accuracy and reliability. Semi-Supervised Domain Adaptation (SSDA) addresses this by leveraging labeled data from some sour...

Manos Giannopoulos, Yi Shen, Michael M. Zavlanos · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.