A remote sensing image landslide segmentation network (VFM-MoME) jointly guided by a vision foundation model and a mixture of Mamba experts is proposed, mitigating the issue of insufficient generalized features in specific landslide study areas and enhancing the model’s ability to handle ambiguous features.
Abstract
The linear computational complexity embraced by Mamba has demonstrated significant application potential in context modeling for the landslide segmentation tasks from remote sensing images. However, existing methods show deficiencies in terms of discrimination and generalization when applied to extreme remote sensing landslide scenarios, such as low resolution and abnormal lighting. To confront these challenges, we propose a remote sensing image landslide segmentation network (VFM-MoME) jointly guided by a vision foundation model and a mixture of Mamba experts. Specifically, we first design a dual-branch joint encoding architecture that integrates a frequency-aware wavelet block as the main encoding branch with the visual foundation model fusion as the auxiliary branch, thereby mitigating the issue of insufficient generalized features in specific landslide study areas. We also construct a mixture of Mamba expert block to enable the decoder to process both global context and local fine-grained features of landslides, addressing the shortcoming of simple serial Mamba in capturing local details and balancing between global semantic relationships and the edges and textural details of objects. Furthermore, we bring in a binary uncertainty enhancement module to guide the model in exploring challenging samples, thus enhancing the model’s ability to handle ambiguous features. Test results on the publicly available datasets of Landslide4Sense and GVLM demonstrate that our method achieves competitive performance.
Semantic segmentation of remote sensing imagery has been widely applied in landslide identification, effectively addressing the time-consuming and labor-intensive nature of manual visual interpretation. However, existing models still face challenges in extracting multiscale features and accurately delineating boundaries under complex background conditions. To overcome these limitations, this study proposes a multiscale boundary aware network (MBANet) for landslide identification in remote sensing imagery. Specifically, we design a multiscale cross-interaction convolution (MCC) module that captures local details and broader contextual cues through heterogeneous receptive-field branches, and recalibrates the concatenated multiscale features via an adaptive cross-branch interaction strategy. In addition, a boundary sensitive refinement attention (BSRA) module is introduced to enhance boundary localization by combining a Sobel-based gradient prior, learnable boundary estimation, and region-context enhancement for fine-grained boundary refinement. These modules are integrated into an encoder–decoder architecture to jointly achieve semantic consistency and boundary precision. Experimental results on two public datasets show that MBANet outperforms other comparison models in overall segmentation performance and maintains competitive performance in boundary delineation. On the Bijie dataset, it achieves a recall of 83.84% and an $F1$ -score of 85.39%; on the Palu dataset, it reaches a recall of 75.83% and an $F1$ -score of 78.17%, highlighting its superior performance.
Zixun Xie, Chuang Song, Xingmin Cai et al.· IEEE Geoscience and Remote S...· 0 citations
The semantic interpretation of remote sensing imagery through segmentation has become indispensable for a wide range of applications, including resource exploration, environmental assessment, and land-use analysis. Yet, accurate parsing of such images remains challenging because complex object boundaries and large scale differences often weaken the ability of conventional Convolutional Neural Network (CNN)-based methods to preserve local details. In response, this study constructs a segmentation framework that couples wavelet convolution with the Mamba architecture. To strengthen feature learning in the intermediate stages, an Auxiliary Segmentation Module (ASM) is employed to provide additional supervisory guidance, which supports optimization and encourages the representation of subtle semantic details. Wavelet-transform convolution is also introduced into the downsampling path, enabling spatial cues and frequency-related information to be exploited in a more coordinated manner for finer boundary and texture modeling. Experiments on public remote sensing datasets and mining area imagery further confirm the effectiveness of the method. Compared with several existing segmentation approaches, the proposed model delivers better overall performance in mIoU, F1-score, and recognition accuracy, particularly in scenes where multiple land-cover categories are heavily interlaced. Moreover, these gains are obtained with relatively low model complexity, suggesting good potential for practical deployment in land monitoring and ecological management.
Wenxi He, Zongmin Yin, Yulong Yang et al.· Remote Sensing· 0 citations
Vision foundation models (VFMs) pretrained on large-scale datasets have significantly improved performance in remote sensing semantic segmentation. However, existing methods typically rely on full fine-tuning, which requires updating all model parameters. Instead of updating the full parameter set, parameter-efficient fine-tuning (PEFT) achieves competitive performance by optimizing only a small subset of parameters. Despite its success, most existing PEFT methods are mainly designed for natural image tasks and fail to account for the unique multiscale characteristics of remote sensing images. To address these challenges, we propose multi-scale cognitive feature refinement (MsRE) tuning, a novel PEFT method tailored for remote sensing semantic segmentation. In particular, MsRE captures multiscale contextual information by applying cognitive operations with different cognitive fields to intermediate features of the backbone. It then introduces a set of learnable tokens to establish interactions with features at different scales, enabling precise feature refinement and progressive feature propagation across network layers. This mechanism enhances the model’s ability to understand complex remote sensing scenes and improves downstream segmentation performance. With significantly fewer trainable parameters, MsRE provides an efficient yet effective solution for adapting VFMs to remote sensing segmentation tasks. Extensive experiments demonstrate that MsRE achieves competitive segmentation performance with substantially fewer trainable backbone parameters, providing a favorable balance between accuracy and parameter efficiency. The project is available at http://woldier.top/MsRE
Bin Wang, Shun Lv, Zhi Li et al.· IEEE Transactions on Geoscie...· 0 citations
Frequency and Edge-guided SAM (FE-SAM) is proposed, a scalable and efficient framework for RSISS that adaptively decomposes and modulates frequency-domain features based on the input data and designs EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image.
Feng Gao, Zizhe Pan, Haoting Wang et al.· IEEE Transactions on Geoscie...· 0 citations
To address the challenges of diverse landslide morphologies and strong background interference in post-earthquake mountainous regions, where a single model suffers from limited adaptability, this paper proposes a dual-model ensemble method for remote sensing landslide identification based on Swin Transformer. The method employs Swin Transformer as a shared backbone network to reduce computational redundancy, integrates the core enhancement modules from SCPD-Deeplabv3+ and LSMFormer to simultaneously capture the global structure of complex landslides and fine-scale boundary details of small landslides, and reuses the MSAD decoder for deep feature fusion to achieve pixel-level segmentation. Experimental results show that the ensemble model achieves a mean intersection over union (mIoU) of 91.88%, with precision, recall, and F1-score of 94.37%, 96.11%, and 94.78%, respectively. The proposed method outperforms each individual model, effectively reducing false positives and missed detections, while balancing the contour accuracy of large-scale landslides and the detail precision of small-scale landslides.
Xiaoyu Fan, Xiaobin Li, T. Ma et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.