A Hybrid Visual Mamba Network With Global-Local Perception for Land Use and Land Cover Semantic Segmentation
Aiming to address limitations in modeling global context and local details in current remote sensing image (RSI) segmentation for land use and land cover, a hybrid visual Mamba model (HVMM) based on global sensing and local details is introduced to gain a deeper understanding of RSIs. The proposed HVMM framework combines two core components: a hybrid visual Mamba block with global and local features (HVMB) and a cross-scale attention aggregation (CSAA) module. The HVMB captures long-term dependency and aggregates multiscale features, thereby boosting the high-fidelity representations. Concurrently, the CSAA module mitigates feature distribution discrepancies across different scales, ensuring robust generalization capabilities. Experiments on benchmark datasets, including LoveDA and UAVid, conclude that HVMM achieves an OA of 77.86% and 88.02%, an mF1 of 80.05% and 80.41, and an mIoU of 63.01% and 71.13%, respectively, notably improving segmentation accuracy and supervision efficiency of land use.