Jul 2026· Scientific Journal of Technology· 0 citations· 37 references
TL;DR
Current evidence suggests that foundation models are most promising when they are deployed as interactive, auditable components within human-in-the-loop workflows, where they can reduce annotation burden, support rapid draft segmentation, and improve consistency across large imaging studies.
Abstract
Medical image segmentation is moving from task-specific convolutional models toward foundation models that can be adapted across organs, modalities, and clinical tasks with fewer manual labels. This transition has been accelerated by self-supervised pretraining, vision-language learning, and promptable segmentation frameworks such as the Segment Anything Model and its medical derivatives. However, the clinical value of these systems cannot be inferred from technical novelty alone. Medical images differ from natural images in dimensionality, intensity statistics, acquisition protocols, disease prevalence, and safety requirements, and recent evaluations show that naive zero-shot transfer remains inconsistent across modalities and lesion types. This narrative review synthesizes literature published up to May 22, 2026, on foundation models for medical image segmentation, with emphasis on technical evolution, application scenarios, validation strategies, and governance needs. Current evidence suggests that foundation models are most promising when they are deployed as interactive, auditable components within human-in-the-loop workflows, where they can reduce annotation burden, support rapid draft segmentation, and improve consistency across large imaging studies. Their translation into routine practice requires external validation, uncertainty-aware quality control, prospective workflow evaluation, bias assessment, and transparent reporting under medical AI guidelines. Future work should prioritize patient-level multimodal modeling, 3D and longitudinal segmentation, federated evaluation, and clinically meaningful endpoints rather than isolated benchmark gains.
Few-shot medical image segmentation (FSMIS) seeks to delineate unseen structures from a small support set, but its standard formulation fixes task-defining evidence before inference. This assumption is fragile under acquisition shift, atypical pathology, ambiguous boundaries, and poor image quality. Adding clinician interaction and rapid adaptation is not sufficient: the binding constraint is deciding when asking or changing is warranted. We therefore reframe FSMIS as a three-layer sequential decision problem. First, decidable self-assessment separates errors that a bounded intervention can repair from those that no admissible intervention can reach. We formalize this distinction through a correctable set defined by the update operator and remaining interaction budget. Second, selective interaction allocates a distinct expert-attention budget by response-conditioned net expected value of information, yielding explicit accept, query, and defer actions. Third, bounded adaptation emphasizes reversibility and independent safety reassessment rather than speed. A complementary cross-case memory stores reproducible correction priors over failure modes instead of disease-specific mask priors. This structure links sparse support representation, cross-domain robustness, multi-level risk estimation, clinician feedback, and governed experience transfer. We state six hypotheses with an explicit dependency order and propose a minimal pilot that can falsify the foundational self-assessment claim before a clinician study. The central claim is not that interaction resolves domain shift, but that scarce expert attention should be used only when a bounded intervention is expected to reach a clinically better outcome.
Recent advancements in foundation models (FMs) have catalyzed a paradigm shift in medical image analysis. Unlike traditional task-specific artificial intelligence (AI) models, FMs leverage large-scale datasets to learn generalized representations that can be adapted to downstream clinical applications. Despite the rapid proliferation of FM research in medical imaging, there is a lack of unified synthesis that systematically maps the evolution of architectures, training paradigms, and clinical applications across modalities. To address this gap, this review provides a comprehensive and structured synthesis of FMs in medical image analysis by systematically organizing studies into two primary categories: vision-only foundation models (VFMs) and vision-language foundation models (VLFMs), based on their architectural foundations, training strategies, and downstream clinical tasks. A quantitative analysis was conducted on both VFMs and VLFMs to characterize temporal trends in dataset utilization and application domains, along with pooled performance and subgroup analyses. We also critically discuss persistent challenges, including cross-domain generalization, computational scalability, FM evaluation, fairness, and deployment. Finally, we identify key future research directions aimed at enhancing the robustness, interpretability, and clinical integration of FMs, thereby accelerating their translation into real-world medical practice.
P. Rajendran, M. Safari, Wen-Feng He et al.· Medical Image Analysis· 0 citations
Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progress, existing FS-MIS solutions span diverse paradigms but are evaluated under inconsistent settings, leaving their relative effectiveness unclear. We introduce FAME, a unified benchmark for evaluating FS-MIS solutions, covering specialists, SAM-based methods, CLIP-based methods, and MLLM-based methods. FAME contains 14,958 test samples across 7 anatomical sites, 9 imaging modalities, and 14 ROI categories, and evaluates models under zero-shot and ten-shot settings with additional assessment of target-absence recognition and generalization under covariate and semantic shifts. Our evaluation reveals several findings. First, effective few-shot segmentation depends on how models exploit support examples: direct visual adaptation generally outperforms prompt-based strategies. Second, increasing support examples improves performance only when models can effectively utilize them. Third, semantic transfer remains substantially more challenging than imaging-domain adaptation, and strong localization ability does not necessarily imply reliable target-absence recognition. We hope FAME provides a comprehensive understanding of current FS-MIS solutions and facilitates the development of more effective and reliable few-shot medical segmentation methods.
Jinghong Liu, Yuchuan Deng, Fanping Liu et al.· arXiv.org· 0 citations
Medical image segmentation is a key component of computer-aided diagnosis and treatment planning. Despite substantial progress in deep learning–based models, most existing approaches depend heavily on large annotated datasets and often fail to generalize across heterogeneous clinical environments, limiting their deployment in real-world settings characterized by domain shifts and scarce expert annotations. This paper presents a zero-shot learning framework named GroundMed-SAM for medical image segmentation. The framework integrates GroundingDINO for prompt-based region localization and MedSAM for mask generation. To address the weak alignment between visual features and medical semantics in GroundingDINO, which is pretrained on general domain image-text pairs, we introduce learnable medical text embeddings that explicitly parameterize domain-specific terminology in a continuous semantic space. These embeddings are optimized during training to better align medical concepts with visual representations, thereby strengthening text-image correspondence and improving detection-guided segmentation. The proposed framework preserves true zero-shot capability, enabling segmentation of previously unseen anatomical structures without task-specific labels. Extensive experiments on multiple public datasets across diverse modalities and clinical contexts demonstrate that our method achieves competitive segmentation performance in-domain while exhibiting superior robustness under cross-domain evaluation. Although supervised baselines outperform the proposed framework by only 3–5% on in-domain datasets, they experience substantial performance degradation when evaluated on unseen domains. Additionally, the framework achieves an AUC of 98.9 in endoscopic polyp detection, highlighting the effectiveness of the proposed medical-aware textual embeddings in guiding region localization. These results demonstrate the effectiveness of the proposed framework in improving cross-domain generalization for medical image segmentation with limited annotations.
V. Nguyen, Hoang Quan Luong, Phuc Ngoc Pham· IEEE International Conferenc...· 0 citations
MRD-UNet provides a practical balance between segmentation accuracy and computational efficiency and outperforms baseline CNNs and performs comparably to heavier transformer-based models while using significantly fewer parameters.
Musa Doğan, I. Ozkan· BMC Medical Imaging· 0 citations
A comprehensive survey of UQ techniques in medical image segmentation is presented, categorizing existing approaches into Bayesian methods, deep ensembles, deterministic methods, test-time data augmentation, and hybrid models, while treating foundation-model-based UQ as a separate cross-cutting category.
Seyed Sina Ziaee, K. Ovens· Journal of Imaging· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.