Bridging Global Context and Local Detail in State Space Models for Fine-Grained Landslide Segmentation in Remote Sensing Imagery
Abstract
Accurate landslide mapping from high-resolution remote sensing imagery requires both broad spatial context and precise boundary detail. State space models (SSMs), such as vision state space duality (VSSD), provide global receptive fields with linear complexity, but their global interaction mechanism offers no dedicated support for the fine local structures required by accurate landslide delineation. We propose landslide state-space Mamba (LSMamba), a landslide segmentation network that augments a VSSD backbone with two lightweight local enhancement modules: a structural feature calibration (SFC) module that adaptively calibrates spatial structural discrepancies after global interaction, and a local detail enhancement module (LDEM) that reinforces multiscale textures before spatial downsampling. Both modules are applied only at the shallow encoder stages where spatial detail is richest, introducing limited incremental computational overhead. For decoding, we design an efficient VSSD-based decoder that replaces heavy convolutional fusion heads, improving segmentation accuracy while substantially reducing computational cost. On three diverse benchmarks—Bijie, globally distributed coseismic landslide dataset, and Landslide4Sense—LSMamba consistently achieves the highest mIoU among all compared methods while maintaining a favorable accuracy–efficiency tradeoff. Overall, the results show that lightweight local enhancement is an effective and efficient strategy for fine-grained landslide mapping with SSM backbones.