Accurate colonoscopic polyp segmentation is essential for early colorectal cancer detection, yet remains challenging due to low contrast, specular highlights, blurred boundaries, and large variations in polyp appearance. We propose DFANet, a dual-branch feature aggregation framework that combines transformer-driven global semantic modeling MiT-B5 with CNN-based fine-grained representation learning EfficientNet-B5. A channel-spatial fusion module adaptively integrates multi-scale encoder outputs through element-wise aggregation, channel attention, and spatial attention, followed by a DSC Transformer that enhances contextual reasoning before decoding. The multi-stage decoder progressively restores resolution while refining structural details to produce accurate segmentation masks. Extensive experiments on five benchmark datasets demonstrate consistent performance gains, with DFANet achieving improvements of up to 3-7% in Dice and mIoU over state-of-the-art methods on Kvasir-SEG, CVC-ClinicDB, BKAI-IGH, ETIS-LaribPolypDB, and CVC-ColonDB.
Kokkanti Srinija, M. Harshini, Panigrahi Srikanth· International Conference on...· 0 citations
Accurate polyp segmentation is essential for early detection of colorectal cancer, where precise delineation of lesion boundaries directly impacts clinical decision-making. Despite significant progress, existing convolutional methods often struggle to capture global contextual information, while transformerbased and foundation models introduce high computational complexity and require extensive fine-tuning. In this work, we propose LS-PolypSeg, a parameter-efficient framework that leverages a pretrained SAM3 vision encoder with Low-Rank Adaptation (LoRA) for domain-specific learning. By selectively adapting key transformer layers while keeping most of the encoder frozen, the proposed approach preserves rich pretrained representations while significantly reducing training overhead. To complement global feature extraction, a lightweight UNetstyle decoder performs multi-scale feature fusion, enabling accurate recovery of fine-grained spatial details. Extensive experiments on three benchmark datasets, namely Kvasir-SEG, CVC-ClinicDB, and BKAI-IGH, demonstrate that LS-PolypSeg achieves competitive and state-of-the-art performance across multiple evaluation metrics. These results highlight the effectiveness of combining foundation model representations with efficient adaptation and hierarchical decoding for robust, scalable polyp segmentation.
Sanjana Jhansi Ganji, Panigrahi Srikanth, Kaushal Sambanna et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.