Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
Abstract
Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
Recent advancements in foundation models (FMs) have catalyzed a paradigm shift in medical image analysis. Unlike traditional task-specific artificial intelligence (AI) models, FMs leverage large-scale datasets to learn generalized representations that can be adapted to downstream clinical applications. Despite the rapid proliferation of FM research in medical imaging, there is a lack of unified synthesis that systematically maps the evolution of architectures, training paradigms, and clinical applications across modalities. To address this gap, this review provides a comprehensive and structured synthesis of FMs in medical image analysis by systematically organizing studies into two primary categories: vision-only foundation models (VFMs) and vision-language foundation models (VLFMs), based on their architectural foundations, training strategies, and downstream clinical tasks. A quantitative analysis was conducted on both VFMs and VLFMs to characterize temporal trends in dataset utilization and application domains, along with pooled performance and subgroup analyses. We also critically discuss persistent challenges, including cross-domain generalization, computational scalability, FM evaluation, fairness, and deployment. Finally, we identify key future research directions aimed at enhancing the robustness, interpretability, and clinical integration of FMs, thereby accelerating their translation into real-world medical practice.
P. Rajendran, M. Safari, Wen-Feng He et al.· Medical Image Analysis· 0 citations
As artificial intelligence becomes increasingly integrated into medical imaging practice, its robustness across heterogeneous real-world settings remains a major challenge. We quantified the effect of real-world distribution shifts on three-dimensional AI models for lung nodule analysis on CT and, motivated by these shifts, developed and evaluated MedStyle-3DG, an open-source 3D vision-language framework for domain generalization. This retrospective multicenter study assembled 2679 chest CTs acquired from 2010 to 2025, with institutional review approval and waiver of informed consent. The study was designed in two stages: first, to quantify the extent of performance degradation from in-distribution (ID) to out-of-distribution (OOD), under clinically realistic distribution shifts; and second, to test whether a domain generalization strategy can mitigate these gaps. Data were split into training, validation, test-ID, and test-OOD to model three shifts: exposure variations, device manufacturer, and geographic changes. We then proposed MedStyle-3DG, a 3D vision-language domain generalization framework combining feature statistics mixing, stochastic weight averaging, vision-language alignment, and a three-branch ensemble. OOD F1 score and the gaps between test-ID and test-OOD were the primary outcomes. The 3D ResNet50 baseline showed consistent ID-OOD degradation across all shifts, largest for exposure (F1, 0.630 vs 0.521; gap, 10.9 percentage points; P < .001). MedStyle-3DG improved exposure OOD F1 to 0.601 and reduced the gap to 7.7 percentage points, with similar improvements for manufacturer and geographic shifts. Across all three factors, MedStyle-3DG achieved state-of-the-art performance in mean OOD F1 (P < .001) and reduction of the F1 generalization gap (P < .041). Real-world protocol, vendor, and population shifts substantially degrade volumetric lung nodule CAD performance. MedStyle-3DG reduces generalization gaps and its code is available at https://github.com/RafaelMedelean/MedCLIP-3DG .
B. Bercean, Rafael Medelean, A. Tenescu et al.· Journal of imaging informati...· 0 citations
Current evidence suggests that foundation models are most promising when they are deployed as interactive, auditable components within human-in-the-loop workflows, where they can reduce annotation burden, support rapid draft segmentation, and improve consistency across large imaging studies.
Juntao Wei· Scientific Journal of Techno...· 0 citations
The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks.
Sam Ansari, Nastaran Faraji, Luke K. Topham et al.· Frontiers in Artificial Inte...· 0 citations
AI-enhanced neuroimaging is rapidly developing, with applications in tumor segmentation, hemorrhage detection, vascular lesion mapping, tractography, and augmented reality. Despite high algorithmic performance in experimental studies, few of these tools are integrated or validated into routine neurosurgeon practice. This review addresses the translational gap by assessing evidence for clinical relevance, explainability, and workflow integration, and by proposing a pragmatic framework implementation for decision support. A narrative review was conducted of literature published between 2015 and 2025 using PubMed, Google Scholar, and regulatory reports. Included studies comprised primary studies, technical reports, and reviews on AI-driven neuroimaging tools with specific pertinence to neurosurgical diagnosis, planning, or intraoperative decision making. Extracted data was synthesized thematically across domains of technical validation and clinical benefit, explainability, and implementation feasibility. Existing AI tools demonstrate promising technical performance: tumor segmentation models achieve Dice scores > 0.80, hemorrhage detection networks report sensitivities over 90% and AUCs ~ 0.95, tractography augmentation and intraoperative image registration enhance mapping accuracy, and AR overlays are increasingly feasible. Nevertheless, the majority of investigations remain retrospective, single-center studies without external validation. Explainability adoption remains variable across studies. Latency, interface design, regulatory uncertainty, and interoperability complicate the workflow integration. And few devices have undergone prospective multicenter trials or secured full regulatory approval. While AI neuroimaging has clear potential to facilitate the accuracy and efficiency of current neurosurgical approaches, translation into the clinical setting necessitates robust multicenter validation, transparent explainability features, and integration into existing surgical workflows. A structured framework connecting validation, explainability, and implementation can assist neurosurgeons and institutions in determining which tools are truly ready for safe adoption from image to incision.