Skip to content
Review Open access

A Review of Deep Learning Methods for Multimodal Medical Image Fusion

Jul 2026 · Italian National Conference on Sensors · Vol 26 · 0 citations · 93 references
Medicine

Abstract

Multimodal medical image fusion (MMIF) aims to integrate complementary information from different imaging modalities into a single, more informative image to support clinical diagnosis. Since 2017, a growing number of deep learning-based approaches have been proposed for MMIF, including various network architectures designed to enhance visual quality. However, a comprehensive and up-to-date review of deep learning-based MMIF techniques is still lacking. To fill this gap, this paper provides a comprehensive survey of deep learning-based MMIF methods. First, we categorize MMIF approaches based on deep learning frameworks and conduct an in-depth analysis of loss functions, evaluation metrics, medical imaging modality pairs, and medical datasets. Then, we review the mainstream deep learning-based MMIF methods, including CNNs, Autoencoders, GANs, and Transformers. Subsequently, we introduce emerging deep learning methods, including diffusion models and Mamba-based methods. In addition, we summarize widely used datasets and evaluation metrics, conduct quantitative experiments on representative methods, and propose a unified set of evaluation metrics for standardized comparison. Finally, we identify key research challenges and outline promising future directions for deep learning-based MMIF. This survey aims to provide researchers with a clear understanding of recent progress in deep learning-based MMIF and to facilitate further studies in this area.

Read PDF

Similar papers

Review Open access Aug 2026

A Comprehensive Review of Multimodal Medical Image Fusion: Techniques, Evaluation, and Future Directions

In recent years, multimodal medical image fusion (MMIF) has attracted significant attention due to its ability to integrate complementary information from different medical imaging modalities and provide more comprehensive information for clinical analysis. By combining anatomical information from modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) with functional information from positron emission tomography (PET) and single-photon emission computed tomography (SPECT), MMIF can improve image quality and support subsequent tasks such as disease analysis, lesion detection, image segmentation, and treatment planning. This review provides a comprehensive overview of MMIF from theoretical and technical perspectives. First, commonly used medical imaging modalities and publicly available medical image databases are summarized and compared. Subsequently, the general workflow, fusion levels, and quality requirements of MMIF are introduced. Representative fusion techniques are then systematically reviewed, including spatial-domain methods, transform-domain methods, sparse representation-based methods, deep learning-based methods, hybrid methods, and emerging Mamba-based approaches. In addition, commonly used image fusion quality assessment metrics are analyzed, and the reported quantitative performance of representative MMIF methods is compared and discussed. Finally, current challenges and future development trends of MMIF are presented, including robustness, clinical translation, and emerging multimodal learning paradigms.

Pengquan Han, Cui-Yin Liu, Cuiwei Wang · 0 citations
Review 2026

A Systematic Literature Review of Deep Learning Models for Medical Image Analysis

It is concluded that while deep learning has achieved remarkable performance in many medical imaging benchmarks, substantial work remains to ensure generalization, interpretability, and ethical deployment in real-world clinicalsettings.

Govinda Sahu · 0 citations
Review Jul 2026

Exploring Deep Transfer Learning for Medical Image Processing and Analysis: A Comprehensive Analysis across Modalities.

Empirical evidence from recent studies demonstrates that fine-tuning and network-based DTL strategies, including federated learning, consistently enhance diagnostic accuracy, robustness, and generalization across multiple medical imaging modalities, particularly in data-limited clinical scenarios.

M. A. S. Banu, A. Dhavapandiammal, K. Palanisamy · 0 citations
#explainable ai Review Open access Aug 2026

Deep Learning in Medical Imaging: Architectures, Clinical Applications, and Emerging Directions

Major deep learning architectures, including CNNs, residual networks, UNet, attention-based models, Vision Transformers, and hybrid approaches, along with their clinical applications are summarized and emerging directions such as self-supervised learning, Explainable AI, federated learning, and lightweight models are highlighted as promising approaches for more reliable and accessible medical image analysis.

Lakshmi Sai Anusha Dadi, Pravallika Devi Kommana · 0 citations
Conference Jul 2026

AI Driven Disease-Specific Adaptive Multimodal Medical Image Fusion with CNN

Despite the utilization of multimodal medical imaging as supplementary anatomical and functional data crucial for precise illness diagnosis, the appropriate integration of multimodal pictures has been challenging due to discrepancies in resolution, contrast, and disease-specific imaging characteristics. Most established techniques for image fusion are modality-specific, depend on manually crafted features, and lack the adaptability to encompass a wide variety of pathological presentations, hence constraining their clinical applicability. This paper proposes an AI-driven, disease-specific, adaptive multimodal medical image fusion utilizing Convolutional Neural Networks (CNNs) to tackle these challenges. This strategy is proposed to be an end-to-end trained modality-specific and pathology-aware feature capable of generating adaptive fusion to increase clinically significant areas, excluding structural characteristics. The qualitative and quantitative assessments reveal that the experimental results of multimodal medical image datasets indicate that the proposed method surpasses both traditional and sophisticated deep learning-based fusion techniques. The performance metrics of PSNR, SSIM, entropy, and mutual information, which indicate enhanced fusion quality, contrast, and diagnostic clarity, substantiate the effectiveness of the proposed framework in facilitating disease-oriented clinical decision-making and computer-aided diagnostics systems.

K. Jameema, T. Sunitha, Maruturi.Haribabu et al. · 0 citations
Conference Jul 2026

Multi-Modal Medical Image Fusion Using Hybrid CNN-Transformer Models for Early Detection of Chronic Diseases

Chronic disease early and accurate detection is a major healthcare issue nowadays, and most importantly, there is the rising prevalence or use of heterogeneous medical image data such as CT, MRI, X-ray, and retinal scans. Conventional models of deep learning such as CNNs perform well on spatial aspects of feature extraction but not generally on long-term relations and overall context. In this paper, we introduce a new hybrid deep learning network combining Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to carry out multi-modal medical image fusion to assist in the diagnostic process with better results. The CNN branch gives fine-grain local representation and the Transformer module will predict the global connections between modalities. Empirical tests on publicly available data show that the proposed model is better than single CNN and Transformer models in using the datasets and there are massive increments in precise score, recall, and F1-score on chronic disease diagnosis in the initial stages. The field of research shows how hybrid architecture can be used to combine complementary information about various scanning modalities, which can become the direction of AI-aided decision-making in prevention.

N.R Azhakeshwari, S. Christy, S. Saranya et al. · 0 citations