Skip to content
Review Open access

Transformers for 3D medical image analysis: a systematic review of architectural innovations, performance, and clinical applications

Jul 2026 · Artificial Intelligence Review · 0 citations

TL;DR

This systematic review systematically evaluates 138 peer-reviewed studies published between January 2017 and December 2025 using the PRISMA 2020 framework to reveal a clear dominance of hybrid CNN–Transformer architectures, with MRI representing the most extensively studied modality.

Abstract

The growing integration of Transformer-based architectures into 3D medical image analysis has driven significant advances across segmentation, classification, detection, registration, and reconstruction tasks. However, existing reviews remain fragmented, often focusing on 2D medical image analysis or specific modalities or tasks without providing a comprehensive, structured synthesis of architectural innovations, benchmark performance, and clinical applicability. This systematic review addresses these gaps by following the PRISMA 2020 framework to systematically evaluate 138 peer-reviewed studies published between January 2017 and December 2025, identified across Scopus, Web of Science, and PubMed. We evaluate Transformer-based and hybrid CNN–Transformer architectures across major 3D imaging modalities; MRI, CT, PET, and ultrasound using a structured five-question research framework addressing architectural evolution, benchmark performance, modality-specific trends, methodological rigour, and reproducibility. To quantify methodological progress beyond performance metrics, we introduce an Architectural Innovation Score grounded in a component-level innovation encoding scheme. Our analysis reveals a clear dominance of hybrid CNN–Transformer architectures, with MRI representing the most extensively studied modality. Multimodal imaging achieves the highest normalized mean performance, followed by MRI, CT, and ultrasound. Emerging paradigms, including State Space Models and diffusion-based Transformers, show promise but remain underexplored. Despite strong benchmark results, critical limitations persist, including insufficient external validation, limited code availability, and inconsistent reporting practices.

Read PDF

Similar papers

Review Open access Aug 2026

A survey of transformer-based architectures in medical image analysis: models, applications, and challenges

The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks.

Sam Ansari, Nastaran Faraji, Luke K. Topham et al. · 0 citations
Review Open access Sep 2026

A Clinically Grounded Review of Medical Image Classification: Quantitative Insights into CNNs, Vision Transformers, and Hybrid CNN-ViT Models.

Medical image classification has advanced substantially with convolutional neural networks (CNNs), Vision Transformers (ViTs), and hybrid CNN-ViT architectures, yet clinical translation remains limited by dataset dependency, inconsistent evaluation practices, and insufficient external validation. This review provides a clinically grounded comparative synthesis of these model families across diverse medical imaging modalities. A structured literature review was conducted across PubMed, IEEE Xplore, Scopus, and Web of Science for studies published between 2016 and 2025. Following predefined eligibility criteria, 81 studies were included in the qualitative review, of which 74 contributed to a dataset-aware descriptive quantitative aggregation. The quantitative synthesis was therefore restricted to descriptive aggregation; a formal meta-analysis was not performed because of substantial methodological heterogeneity and insufficient reporting of study-level variance information across the included studies. CNN-based models demonstrated the most consistent performance, achieving the highest weighted accuracy (0.934) and weighted recall (0.901). ViT-based models achieved competitive weighted accuracy (0.906) and recall (0.893), particularly for OCT and X-ray imaging, but appeared more sensitive to dataset scale and quality. Hybrid CNN-ViT models achieved a weighted accuracy of 0.814 and weighted recall of 0.698, with the greatest performance variability. Only 10 of the 81 reviewed studies (12.3%) reported independent external validation, while calibration and other clinically relevant evaluation measures were inconsistently reported. CNNs provide a robust baseline for medical image classification, whereas ViT- and hybrid-based architectures offer complementary strengths under appropriate data and training conditions. However, limited external validation and inconsistent reporting indicate that strong retrospective performance should not be interpreted as evidence of clinical readiness. Future research should prioritise externally validated, interpretable, and clinically deployable AI systems supported by standardised evaluation practices.

H. Hussaini, Shahana Bano, E. Elyan et al. · 0 citations
Review Aug 2026

Foundation models in medical image analysis: A systematic review and quantitative analysis.

Recent advancements in foundation models (FMs) have catalyzed a paradigm shift in medical image analysis. Unlike traditional task-specific artificial intelligence (AI) models, FMs leverage large-scale datasets to learn generalized representations that can be adapted to downstream clinical applications. Despite the rapid proliferation of FM research in medical imaging, there is a lack of unified synthesis that systematically maps the evolution of architectures, training paradigms, and clinical applications across modalities. To address this gap, this review provides a comprehensive and structured synthesis of FMs in medical image analysis by systematically organizing studies into two primary categories: vision-only foundation models (VFMs) and vision-language foundation models (VLFMs), based on their architectural foundations, training strategies, and downstream clinical tasks. A quantitative analysis was conducted on both VFMs and VLFMs to characterize temporal trends in dataset utilization and application domains, along with pooled performance and subgroup analyses. We also critically discuss persistent challenges, including cross-domain generalization, computational scalability, FM evaluation, fairness, and deployment. Finally, we identify key future research directions aimed at enhancing the robustness, interpretability, and clinical integration of FMs, thereby accelerating their translation into real-world medical practice.

P. Rajendran, M. Safari, Wen-Feng He et al. · 0 citations
Open access Aug 2026

Attention-Enhanced Bimodal 3D Medical Image Segmentation with Two-Stage Learning

Computer-aided diagnostic technologies have demonstrated substantial advantages in 3D medical image segmentation, particularly in multimodal 3D medical image segmentation tasks, where they play a pivotal role in driving continuous innovation in related architectures. As an integration of U-Net and Transformer, the UNETR architecture has demonstrated remarkable efficacy in 3D medical image segmentation. Nevertheless, despite its successes, UNETR remains challenged by clinical complexities such as intricate tumor localization and anatomical structural diversity in complex clinical settings. To address these issues, we propose an enhanced 3D segmentation framework, UAtten-Unetr, designed to improve segmentation accuracy and robustness in complex medical scenarios. The framework captures global contextual information via hierarchical Transformer layers and incorporates a spatial–channel attention module to enable adaptive fusion of multimodal features, thereby effectively enhancing cross-modal feature alignment capabilities. Concurrently, we innovatively developed a unified loss function based on bimodal modality-specific Dice constraints and uncertainty regularization, optimized for synchronous learning across the ACDC (cardiac MRI) and AMOS22 (abdominal CT/MRI) datasets. Experimental results showed that UAtten-Unetr achieved an average Dice score of 92.20% on the ACDC dataset, exceeding the reported nnU-Net result of 91.61% by 0.59 percentage points. On the AMOS22 dataset, the proposed method achieved an average Dice score of 84.51%, exceeding the reported UNETR result of 78.33% by 6.18 percentage points. However, its myocardium Dice score (84.11%) was lower than those of nnU-Net (89.24%) and MT-UNet (89.04%), indicating a remaining limitation in myocardium boundary segmentation. These results indicate competitive segmentation performance under the reported experimental settings. This method delivers dual improvements in accuracy and generalization across complex anatomical scenarios, providing an effective solution for precise diagnosis in intricate clinical environments.

Mengxuan Li, Hao-Yu Wang · 0 citations
Review Open access Aug 2026

Evaluation metrics for synthetic medical imaging.

BACKGROUND Advances in generative artificial intelligence (AI) have accelerated the development and application of synthetic medical imaging. Despite this rapid progress, the evaluation of synthetic medical images remains heterogeneous, with numerous metrics proposed to assess fidelity, realism, diversity, and clinical validity. Currently, no standardized framework exists to guide the selection, interpretation, or comparison of these metrics, limiting reproducibility and cross-study comparability. This systematic review aims to comprehensively summarize and categorize existing metrics used to assess these complementary dimensions of synthetic medical images. METHODS A systematic review was conducted in accordance with PRISMA guidelines. PubMed/MEDLINE, EMBASE, Scopus, and arXiv were searched for studies published between 2015 and April 30, 2025, supplemented by citation screening of included studies. Eligible studies were full-text articles that applied or proposed metrics to evaluate the fidelity, realism, diversity, and/or clinical validity in synthetic medical images. RESULTS A total of 47 studies were included. Evaluation practices were highly heterogeneous. Expert evaluation (n = 25, 53%) and reference-based evaluations were most common (n = 25, 53%), followed by no-reference metrics (n = 24, 51%), and task-based evaluations (n = 24, 51%). The most commonly used individual metrics were peak signal-to-noise ratio (PSNR) (n = 16, 34%), structural similarity index (SSIM) (n = 15, 32%), mean absolute error (MAE) (n = 12, 26%), and Fréchet Inception Distance (FID) (n = 12, 26%). CONCLUSION Evaluation strategies for synthetic medical imaging showed substantial variability and no single metric captured fidelity, realism, diversity, and clinical validity simultaneously. Metric choice is often dictated by data availability rather than clinical purpose. A task-specific, layered evaluation framework could improve comparability and facilitate clinical adoption.

D. D. de Wilde, Benjamin Schärli, Kym Ackermann et al. · 0 citations
Review Open access Sep 2026

Understanding and Mitigating Distribution Shifts in Volumetric Lung Nodule CAD Using a 3D Vision-Language Framework.

As artificial intelligence becomes increasingly integrated into medical imaging practice, its robustness across heterogeneous real-world settings remains a major challenge. We quantified the effect of real-world distribution shifts on three-dimensional AI models for lung nodule analysis on CT and, motivated by these shifts, developed and evaluated MedStyle-3DG, an open-source 3D vision-language framework for domain generalization. This retrospective multicenter study assembled 2679 chest CTs acquired from 2010 to 2025, with institutional review approval and waiver of informed consent. The study was designed in two stages: first, to quantify the extent of performance degradation from in-distribution (ID) to out-of-distribution (OOD), under clinically realistic distribution shifts; and second, to test whether a domain generalization strategy can mitigate these gaps. Data were split into training, validation, test-ID, and test-OOD to model three shifts: exposure variations, device manufacturer, and geographic changes. We then proposed MedStyle-3DG, a 3D vision-language domain generalization framework combining feature statistics mixing, stochastic weight averaging, vision-language alignment, and a three-branch ensemble. OOD F1 score and the gaps between test-ID and test-OOD were the primary outcomes. The 3D ResNet50 baseline showed consistent ID-OOD degradation across all shifts, largest for exposure (F1, 0.630 vs 0.521; gap, 10.9 percentage points; P < .001). MedStyle-3DG improved exposure OOD F1 to 0.601 and reduced the gap to 7.7 percentage points, with similar improvements for manufacturer and geographic shifts. Across all three factors, MedStyle-3DG achieved state-of-the-art performance in mean OOD F1 (P < .001) and reduction of the F1 generalization gap (P < .041). Real-world protocol, vendor, and population shifts substantially degrade volumetric lung nodule CAD performance. MedStyle-3DG reduces generalization gaps and its code is available at  https://github.com/RafaelMedelean/MedCLIP-3DG .

B. Bercean, Rafael Medelean, A. Tenescu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.