Skip to content
Conference

ResUNet++ -ViT: A Fusion Framework for Accurate High-Resolution Medical Image Segmentation

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1-8 · 0 citations · 15 references

Abstract

Segmentation of medical images is a crucial process for diagnosis and treatment planning. Yet, traditional CNN based models are often not able to convey high-level spatial relationships and intricate tissue boundaries in high-resolution medical images. To tackle these issues, this paper presents a novel ResUNet++–Vision Transformer (ResUNet++–ViT) framework incorporating both multi-scale local feature extraction and global contextual learning. The ResUNet++ backbone consists of residual blocks and nested skip connections, which are used to extract hierarchical features, and the Vision Transformer uses self-attention mechanisms to capture long-range dependencies. A fusion module allows for the integration of local and global features, resulting in better segmentation accuracy and preserving the boundary. The proposed model was tested on the ISBI 2012 Electron Microscopy Segmentation Challenge (EMSC) dataset, and obtained a Dice score of 0.960, IoU of 0.920, precision of 0.967, recall of 0.958, accuracy of 0.984 and Hausdorff distance of 2.76. The proposed framework is compared with FCN, UNet, Attention UNet, ResUNet, UNet++, Vision Transformer and TransUNet, and the results show its superiority. The results show that the combination of ResUNet++ and Vision Transformers greatly enhances the performance of segmentation, boundary delineation, and generalization in the field of advanced medical image analysis applications.

View source

Similar papers

Open access 2026

Hybrid U-Net++–Vision Transformer Fusion for Accurate Brain MRI Image Segmentation and classifications

Automatic brain Magnetic Resonance Imaging (MRI) analysis plays a crucial role in computer-aided diagnosis, treatment planning, and disease monitoring by enabling accurate delineation and identification of brain tumors. However, manual segmentation is labor-intensive, time-consuming, and susceptible to inter-observer variability, making automated and reliable methods highly desirable. This paper proposes a Hybrid U-Net++ and Vision Transformer (ViT) Fusion framework for automatic brain MRI tumor segmentation and segmentation-guided classification. The proposed architecture employs a U-Net++ encoder with nested skip connections to extract hierarchical multi-scale features, while a Vision Transformer captures long-range spatial dependencies and global contextual information through self-attention mechanisms. A multi-stage feature fusion strategy integrates convolutional and transformer representations, complemented by boundary-aware refinement to improve tumor localization and preserve fine anatomical details. The segmentation model achieved a training Dice score of 95.6%, IoU of 91.4%, and validation Dice score of 92.6%, demonstrating stable convergence and strong generalization capability. Furthermore, the proposed framework attained a Dice score of 95.8%, IoU of 92.1%, and Hausdorff Distance of 3.35 mm in the ablation analysis, confirming the effectiveness of multi-stage feature fusion and boundary-aware refinement. The segmented tumor regions are subsequently utilized for Region of Interest (ROI)-based classification using deep feature extraction, Global Average Pooling (GAP), and fully connected layers. The classification module distinguishes glioma, meningioma, pituitary tumor, and no-tumor cases, achieving an overall accuracy of 97.55%, macro-average precision of 97.13%, recall of 97.50%, F1-score of 97.31%, and specificity of 98.40%. Comparative analysis further demonstrates the effectiveness of the proposed hybrid architecture over conventional CNN-based and transformer-based approaches. The combined segmentation and classification results indicate that the proposed framework provides an effective approach for tumor localization and tumor-type identification, with potential applicability in automated brain MRI analysis and clinical decision support.

Malathi Janapati, Shaheda Akthar · 0 citations
Open access Aug 2026

WVM-UNet: A Wavelet–Vision Mamba Framework for Enhanced Medical Image Segmentation

Results indicate the WVM-UNet architecture effectively captures discriminative features for precise medical image segmentation, and demonstrates the competitive performance of the method on multiple public datasets.

Yulong Yang, Wenchao Gao, Zheng-Guo Wu et al. · 0 citations
Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Sep 2026

LMDAU-Net: An Effective Lightweight Multi-scale Deformation Aggregation U-Net for Skin Lesion Segmentation.

Automatic skin lesion segmentation is a pivotal problem in the medical domain and an indispensable component in the computer-aided diagnosis program. Most convolutional neural network-based segmentation algorithms have demonstrated promising performance due to their ability to encode detail and semantic features efficiently. However, they fail to capture the long-range contextual information at the global level. Therefore, researchers employ Transformer architecture to address this issue. Unfortunately, these methods fail to learn sufficient pixel information at the local level. Motivated by this, some researchers attempt to design a hybrid architecture based on CNN and Transformer. However, the large number of parameters and high computational cost make them challenging to train and use. To alleviate these problems, we propose an effective Lightweight Multi-scale Deformation Aggregation U-Net (LMDAU-Net), which consists of a Lightweight Local-global Learning Module (LLM) and an Adaptive Interactive Fusion Module (AIF). Specifically, we utilize the two branches of the proposed LLM to efficiently learn local fine-grained and global coarse-grained features that assist the model in capturing the complementary feature representations. Moreover, we employ the AIF to selectively learn semantic and detail features at different scales, which can dynamically explore variable feature cues. Extensive experiments on four skin benchmarks, including ISIC2016, ISIC 2017, ISIC2018, and PH2, demonstrate that LMDAU-Net achieves state-of-the-art performance in both qualitative and quantitative aspects. We have released our code on https://github.com/Lm0611/LMDAU-Net.

Jun-Han Hu, Ming Liu, Jing Yang et al. · 0 citations
Open access Aug 2026

Integrating state space models and attention mechanisms for brain tumor segmentation in MRI

Brain tumor segmentation from MRI is clinically critical yet challenging due to heterogeneous appearance and irregular boundaries. Conventional CNN based methods lack effective global context modeling, while transformer-based approaches are computationally expensive and unstable on limited datasets. To address these, we propose MMA-UNet, a novel hybrid architecture that integrates multiple complementary mechanisms within an encoder-decoder framework. The model employs an EfficientNet-B5 encoder, MedNeXt bridge blocks, a Mamba-based VSS bottleneck, a CBAM enhanced decoder, and deformable refinement for precise boundary adaptation. The model achieves a Dice score of 0.9063, IoU of 0.8304, precision of 0.9043, recall of 0.9092, and specificity of 0.9982 on FigShare benchmark, outperforming the evaluated U-Net, Attention U-Net, TransUNet and Swin UNet under the adopted experimental protocol and maintaining parameter efficiency (30.41 M), demonstrating strong robustness on T1-weighted contrast-enhanced MRI. The proposed model operates on independent 2D slices without volumetric context, and its generalizability to larger datasets remains to be established.

Saritha Saladi, Riyaz Hussain Shaik, Abhi Chevuri et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.