Skip to content
Open access

Deformable multi-head cross attention-based 3D U-net for context-aware multi-modal cardiac image segmentation with uncertainty-aware maps

Jul 2026 · Scientific Reports · Vol 16 · 0 citations · 33 references
Medicine

TL;DR

This work suggests a deformable Multi-Head Cross-Attention mechanism (MHCA) for effective feature fusion and uses a 3D U-Net for hierarchical feature extraction in a 3D U-Net for hierarchical feature extraction.

Abstract

Highly efficient medical image segmentation is crucial for accurate tumor diagnosis in multi-modal imaging datasets having Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) images. Nevertheless, current algorithms have limitations while incorporating supplementary data from multiple modalities, which affects the segmentation performance and interpretability. This proposed study uses a 3D U-Net for hierarchical feature extraction. The 3D U-Net improves segmentation performance by strengthening the model’s capacity to capture local and global relationships. The Swin transformer used in the proposed work helps in context-aware reasoning. This work suggests a deformable Multi-Head Cross-Attention mechanism (MHCA) for effective feature fusion. The uncertainty-aware map and SHAPely-based explainability help the experts in decision-making. The proposed system has demonstrated good performance during the experimental evaluation with multiple datasets. The proposed method achieves a Dice Coefficient of 91.8% on Medical Decathlon (MD), 92.3% on Automated Cardiac Diagnosis Challenge (ACDC), and 93.2% on Multi-Modality Whole Heart Segmentation (MMWHS) datasets, overcoming the baseline models.

Read PDF

Similar papers

Open access Jul 2026

A 3D residual U-Net with attention-driven spatial pyramid pooling for accurate multimodal MRI brain tumor segmentation

Medical imaging plays a crucial role in the accurate detection and localization of brain tumors, which is essential for effective clinical diagnosis and treatment planning. However, conventional segmentation approaches often struggle to capture complex spatial dependencies in volumetric data. To address this limitation, this study proposes an enhanced 3D U-Net architecture for multi-modal MRI-based brain tumor segmentation. The proposed model leverages three-dimensional convolutional operations to effectively capture contextual and spatial information from volumetric inputs. Additionally, an automated preprocessing pipeline, including image resizing, intensity normalization, and data augmentation, is incorporated to improve model robustness and generalization. The performance of the proposed model is evaluated against a conventional U-Net and a ResNet-based segmentation model using standard metrics such as Dice coefficient, accuracy, Intersection-over-Union (IoU), precision, recall, and F1-score. Experimental results demonstrate that the proposed 3D U-Net achieves superior performance, with a Dice coefficient of 0.83 and a Jaccard index of 0.82, outperforming baseline models across all evaluation metrics. Furthermore, the model exhibits improved convergence behavior and reduced overfitting, indicating strong generalization capability. These findings highlight the effectiveness of the proposed approach for volumetric medical image segmentation. Future work will focus on optimizing hyperparameters, enhancing architectural design, and validating the model on larger and more diverse clinical datasets.

Retinderdeep Singh, C. Prabha, Navita Gupta et al. · 0 citations
Open access Jul 2026

DynU-Net: Dynamic Uncertainty-Aware Multi-task U-Net for Joint Lesion Segmentation and Classification in Medical Imaging

A Dynamic Uncertainty-aware Network (DynU-Net) is proposed, a multi-task framework that adaptively balances segmentation and classification through learnable per-task uncertainty parameters that consistently outperforms both single-task and existing multi-task baselines.

Ngoc Ly Tran, Thi Thu Thuy Nguyen, Ba-Hung Ngo et al. · 0 citations
Open access Aug 2026

Attention-Enhanced Bimodal 3D Medical Image Segmentation with Two-Stage Learning

Computer-aided diagnostic technologies have demonstrated substantial advantages in 3D medical image segmentation, particularly in multimodal 3D medical image segmentation tasks, where they play a pivotal role in driving continuous innovation in related architectures. As an integration of U-Net and Transformer, the UNETR architecture has demonstrated remarkable efficacy in 3D medical image segmentation. Nevertheless, despite its successes, UNETR remains challenged by clinical complexities such as intricate tumor localization and anatomical structural diversity in complex clinical settings. To address these issues, we propose an enhanced 3D segmentation framework, UAtten-Unetr, designed to improve segmentation accuracy and robustness in complex medical scenarios. The framework captures global contextual information via hierarchical Transformer layers and incorporates a spatial–channel attention module to enable adaptive fusion of multimodal features, thereby effectively enhancing cross-modal feature alignment capabilities. Concurrently, we innovatively developed a unified loss function based on bimodal modality-specific Dice constraints and uncertainty regularization, optimized for synchronous learning across the ACDC (cardiac MRI) and AMOS22 (abdominal CT/MRI) datasets. Experimental results showed that UAtten-Unetr achieved an average Dice score of 92.20% on the ACDC dataset, exceeding the reported nnU-Net result of 91.61% by 0.59 percentage points. On the AMOS22 dataset, the proposed method achieved an average Dice score of 84.51%, exceeding the reported UNETR result of 78.33% by 6.18 percentage points. However, its myocardium Dice score (84.11%) was lower than those of nnU-Net (89.24%) and MT-UNet (89.04%), indicating a remaining limitation in myocardium boundary segmentation. These results indicate competitive segmentation performance under the reported experimental settings. This method delivers dual improvements in accuracy and generalization across complex anatomical scenarios, providing an effective solution for precise diagnosis in intricate clinical environments.

Mengxuan Li, Hao-Yu Wang · 0 citations
Aug 2026

CERD3D-UNet: context-enhanced residual 3D U-Net with dual-attention gates and hyperparameter optimization for multimodal brain tumor segmentation

CERD3D-UNet is introduced, a context-enhanced residual-dense 3D U-Net model with dual-attention mechanisms and hyperparameter optimization for accurate delineation of whole tumor, tumor core (TC), and enhanced tumor (ET).

Anusha Kakumanu, Venkatramaphanikumar Sistla, Venkata Krishna Kishore Kolli · 0 citations
Jul 2026

EEA-UNet: An efficient element-wise adaptive attention-based network for abdominal multi-organ segmentation.

Accurate X-ray computed tomography (CT) image segmentation of the abdominal organs is a key task in automated medical image analysis, with crucial applications in clinical decision-making, computer-aided diagnosis, and surgical planning. However, existing methods still face significant challenges: insufficient capability in modeling long-range contextual dependencies, hindering the adaptability to the complicated morphological variations and spatial relationships of abdominal organs; and inaccurate boundary segmentation due to blurred edges and irregular anatomical structures, particularly in regions with high tissue adhesiveness. To address these issues, we propose an efficient abdominal multi-organ segmentation model, EEA-UNet. Specifically, we design an efficient element-wise adaptive (EEA) attention mechanism integrated into the skip connections to enhance inter-organ feature interactions while maintaining computational efficiency. This module effectively expands the receptive field, improving long-range dependency modeling. An enhanced multi-scale feature fusion (EMF) module is introduced to strengthen decoding capability, coupled with an edge-awareness composite loss function to optimize segmentation accuracy for small organs and boundary regions.Experimental results on the Synapse dataset demonstrate the competitive performance of EEA-UNet, achieving a Dice score of 84.45% and an HD95 of 0.16. Our method demonstrates a favorable trade-off between segmentation accuracy and computational efficiency, showing improved results compared with several existing approaches in both visual comparison and quantitative metrics.

Panpan Wu, Runpeng Guo, Ziping Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.