A heterogeneous U-shaped architecture that strategically deploys specialized mechanisms based on component-specific functional requirements to optimize multi-scale feature integration and introduces Spatial-Channel Synergistic Attention in skip connections.
Abstract
Based on U-shaped architecture for medical image segmentation faces three fundamental challenges limiting clinical deployment: (1) architectural homogeneity problems where uniform mechanism deployment overlooks distinct encoder-decoder requirements, (2) skip connection feature fusion limitations with insufficient spatial-channel attention integration, and (3) computational efficiency
vs
. performance trade-offs. We propose a heterogeneous U-shaped architecture that strategically deploys specialized mechanisms based on component-specific functional requirements. Our approach utilizes Vision Mamba with Non-causal State Space Duality (VSSD) in encoder/bottleneck for efficient global context extraction, Bi-Level Routing Attention (BRA) in decoder for adaptive detail recovery, and introduces Spatial-Channel Synergistic Attention (SCSA) in skip connections to optimize multi-scale feature integration with only 0.01M additional parameters. Extensive experiments across four diverse datasets demonstrate exceptional performance: for example on Synapse dataset, our model achieves 84% Dice Similarity Coefficient and 13.76 mm Hausdorff Distance with only 23.43M parameters. Details are available on
https://github.com/Yuyan-Bin/Synergistic-Network
.
This work proposes an uncertainty-aware efficient segmentation framework synergizing Mamba state-space models with evidential deep learning, employing a 2D-adapted selective state-space mechanism to capture long-range dependencies with linear complexity O(L), overcoming transformers' quadratic scaling.
TvaraNet is pro-posed, an extremely lightweight segmentation network designed to preserve boundary fidelity under strict efficiency constraints and achieves competitive or superior boundary-aware performance compared to heavier architectures.
Experimental results demonstrate that LightVM-SparseUNet achieves segmentation competitive with state-of-the-art large-scale models across two authoritative public datasets.
Haojie Fan, Kang Xu, Xiaoyu Hou et al.· Biomedical engineering and p...· 0 citations
Multi-scale feature fusion is a cornerstone of encoder-decoder architectures in medical image segmentation, yet effectively integrating representations across stages remains a significant challenge due to the inherent semantic–spatial gap. Deep features encode abstract semantic context but lack spatial precision, whereas early-stage features preserve fine-grained details but suffer from limited semantic discriminability. Existing fusion mechanisms, which often rely on symmetric aggregation or simple skip connections, fail to explicitly model the semantic-to-spatial guidance necessary for precise alignment. To address this, we propose a Tri-stream Prototype Fusion Network (TSPFusion) that introduces three key innovations: (i) a tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage; (ii) a Global Prototype Bank (GPB) that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing persistent semantic priors across images; and (iii) a Detail–Semantic Feature Aligner (DSFA) that performs semantic-guided refinement of spatial features prior to fusion, preventing feature interference from direct concatenation. Additionally, an Adaptive Pyramid Context Decoder module aggregates multi-scale information with resolution-aware dynamic pooling, and a Gradient-Gated Spatial Attention head enforces boundary-sensitive structural consistency. Extensive experiments on four medical imaging benchmarks (CT and Ultrasound) demonstrate that TSPFusion achieves state-of-the-art performance 97.82±0.85% DSC on COVID19 lung CT, 81.76% DSC on COVID19-Seg, and 88.16% mDice on cross-dataset BUSI→\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow $$\end{document}STU, while maintaining a compact 5.92M parameter footprint.
Mohammed A. M. Elhassan, Qianfa Yuan, Zhizhong Xu et al.· Journal of King Saud Univers...· 0 citations
Accurate 3D medical image segmentation is crucial for clinical use. CNNs and transformers have limitations: limited receptive fields or quadratic complexity. State space models like Mamba offer linear complexity with global perception, but existing methods (e.g., EM-Net) struggle with small organs due to poor spatial dependency encoding. We propose ADSAM-Net, integrating a Dual-stage Attention Module (DSAM) into Mamba layers and decoder stages. DSAM enables cross-dimensional feature calibration and detail enhancement. Experiments on Synapse and BTCV datasets show that ADSAM-Net achieves average Dice scores of $\mathbf{7 8. 5 0 \%}$ and $\mathbf{7 8. 5 8 \%}$, significantly improving small-organ segmentation (e.g., pancreas, gallbladder) while maintaining high performance on large organs. Hausdorff distance is also substantially reduced. ADSAM-Net offers an efficient and accurate solution for 3D medical image segmentation.
Yifei Zhang, Jia-Mi Yang, Guan-Yu Lu et al.· 2026 3rd World Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.