Skip to content

DSGH-Net: Medical Image Segmentation via Dual-Statistical Dynamic Context and Graph-Convolutional Heterogeneous Decoder.

Jul 2026 · Journal of imaging informatics in medicine · 0 citations · 16 references
Medicine

TL;DR

The Dual-Statistical Context Modulation Block (DCM-Block), which integrates Global Average Pooling and Global Max Pooling to generate content-aware dynamic weights, thereby achieving dynamic multi-scale feature fusion at the bottleneck layer is proposed.

View source

Similar papers

Open access Jul 2026

SegRWKV: Fast and Accurate Biomedical Image Segmentation via a Receptance-Weighted Key-Value Network

SegRWKV, a hybrid architecture that integrates pretrained Vision-RWKV modules in the encoder for efficient global dependency modeling, and a PixelRefinement module in the decoder to improve feature reconstruction and multi-scale alignment, is proposed, highlighting SegRWKV’s superior ability to balance accuracy, efficiency, and scalability.

Tianyu Cao, Mingyao Ma, Haoxiang Zhang et al. · 0 citations
Open access Sep 2026

Cross-layer semantic alignment and context enhancement network for medical image segmentation

Medical image segmentation aims to accurately delineate organs, tissues, or lesion regions from complex medical images. However, existing hybrid models based on Transformers and convolutional neural networks still suffer from limitations in local detail modeling and cross-layer feature fusion, which often leads to blurred boundary information and loss of structural details. To address these issues, this paper proposes a Cross-layer Semantic Alignment and Context Enhancement Network for medical image segmentation. Specifically, a semantic enhancement module is introduced into the skip connections to achieve effective fusion of high-level semantic information and shallow spatial details through spatial-channel collaborative modeling and multi-scale context extraction (MCE). In addition, a lightweight boundary refinement mechanism is employed in the decoder stage to improve the recovery capability for complex boundary regions. Experiments conducted on the Synapse, ACDC, and GlaS datasets demonstrate that the proposed method outperforms mainstream approaches in terms of Dice, HD95, and Intersection over Union metrics, validating the effectiveness of the cross-layer semantic alignment mechanism for complex medical image segmentation tasks.

Bing Liu, Xin-Xin Sun, Ge-Yi Zhan · 0 citations
Open access Jul 2026

Seed-net: structure-enhanced encoder-decoder network via dual-attention bridge for 2D medical image segmentation

Medical image segmentation (MIS) plays a crucial role in clinical diagnosis and disease prediction. However, existing encoder-decoder segmentation networks still face several structural limitations. First, conventional downsampling operations may discard high-frequency boundary details, resulting in blurred lesion edges. Second, shallow convolutional features are easily affected by local background noise and scale-variable lesion structures, while deep semantic representations often lack effective long-range dependency modeling, limiting global contextual understanding. Third, highly compressed bottleneck features often mix lesion semantics with background artifacts, which enlarges the semantic gap between the encoder and decoder. Finally, standard decoding operations may introduce upsampling artifacts and fail to accurately reconstruct irregular anatomical boundaries. To address these problems, we propose a structure-enhanced encoder-decoder network, termed SEED-Net, for 2D MIS. Specifically, Haar Wavelet Downsampling is introduced to preserve high-frequency structural information during feature compression, while Deep Mamba layers are deployed in deep semantic stages to capture long-range dependencies with linear computational complexity. In the encoder, the proposed encoder gradient multi-scale module introduces a purification-before-interaction strategy, where Grouped Dynamic Gating first recalibrates group-wise structural responses before local-global multi-scale feature interaction, thereby enhancing lesion-relevant structures while suppressing redundant background activations. At the bottleneck, the Large Kernel Dual-Attention Bridge recalibrates spatial and channel responses to isolate lesion-related semantics from background artifacts. In the decoder, the decoder gradient multi-scale Aggregation module improves boundary reconstruction and suppresses upsampling-induced artifacts through multi-receptive feature aggregation. Experimental results on five public datasets, including DSB2018, ISIC 2016, Kvasir-SEG, DRIVE,and LIDC-IDRI, demonstrate the effectiveness of SEED-Net. Specifically, on ISIC 2016, SEED-Net achieves an IoU of 85.54%, Dice of 91.61%, ACC of 95.64%, Spe of 96.25%, and Sen of 93.21%. These results indicate that SEED-Net can effectively preserve fine structural details, reduce background interference, and produce accurate lesion boundaries, showing its potential for reliable MIS.

Zikai Wang, Biyuan Li, Jinying Ma et al. · 0 citations
Conference Jul 2026

GroundMed-SAM: Prompt-based Zero-shot Medical Image Segmentation

Medical image segmentation is a key component of computer-aided diagnosis and treatment planning. Despite substantial progress in deep learning–based models, most existing approaches depend heavily on large annotated datasets and often fail to generalize across heterogeneous clinical environments, limiting their deployment in real-world settings characterized by domain shifts and scarce expert annotations. This paper presents a zero-shot learning framework named GroundMed-SAM for medical image segmentation. The framework integrates GroundingDINO for prompt-based region localization and MedSAM for mask generation. To address the weak alignment between visual features and medical semantics in GroundingDINO, which is pretrained on general domain image-text pairs, we introduce learnable medical text embeddings that explicitly parameterize domain-specific terminology in a continuous semantic space. These embeddings are optimized during training to better align medical concepts with visual representations, thereby strengthening text-image correspondence and improving detection-guided segmentation. The proposed framework preserves true zero-shot capability, enabling segmentation of previously unseen anatomical structures without task-specific labels. Extensive experiments on multiple public datasets across diverse modalities and clinical contexts demonstrate that our method achieves competitive segmentation performance in-domain while exhibiting superior robustness under cross-domain evaluation. Although supervised baselines outperform the proposed framework by only 3–5% on in-domain datasets, they experience substantial performance degradation when evaluated on unseen domains. Additionally, the framework achieves an AUC of 98.9 in endoscopic polyp detection, highlighting the effectiveness of the proposed medical-aware textual embeddings in guiding region localization. These results demonstrate the effectiveness of the proposed framework in improving cross-domain generalization for medical image segmentation with limited annotations.

V. Nguyen, Hoang Quan Luong, Phuc Ngoc Pham · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.