Jul 2026· Engineering Research Express· Vol 8, pp. 155214· 0 citations· 25 references
Physics
TL;DR
Seasonal, scenario-based, complexity, ablation, and residual error analyses demonstrate that SFG-UNet improves glacier continuity, boundary recovery, and robustness under complex optical imaging conditions while maintaining an acceptable computational cost.
Abstract
Glacier segmentation in optical remote sensing imagery remains challenging in complex mountain environments due to fragmented glacier structures, blurred boundaries, seasonal snow confusion, terrain shadows, bare rock, and cloud interference. To address these issues, this study proposes a state-space-guided U-Net framework, termed SFG-UNet, for glacier segmentation in optical remote sensing imagery. The model introduces a long-range state space block in the encoder to enhance global contextual representation, an SSM-guided frequency decoupling and boundary calibration module in the skip pathway to refine low- and high-frequency features, and a semantic-guided full-scale gated fusion module in the decoder to improve selective multi-scale feature aggregation. Experiments were conducted on a self-built Landsat-8/9 glacier dataset from the Animaqing Snow Mountain region and an independent public DL4GAM Alps dataset for external validation. On the Animaqing dataset, SFG-UNet achieved 95.71% accuracy, 95.04% dice, 90.08% kappa, and 90.61% MIoU, outperforming representative CNN-based, attention-based, Transformer-based, frequency-domain, glacier-oriented, and SSM-based segmentation methods. On the external DL4GAM Alps dataset, SFG-UNet also achieved the best overall performance, with 89.86% accuracy, 89.12% dice, 81.02% kappa, and 83.64% MIoU. Seasonal, scenario-based, complexity, ablation, and residual error analyses further demonstrate that SFG-UNet improves glacier continuity, boundary recovery, and robustness under complex optical imaging conditions while maintaining an acceptable computational cost.
Accurate delineation of landslides in RGB optical remote sensing imagery supports rapid disaster mapping and post-event assessment. This remains difficult because landslides are often small and irregular, resemble bare soil or disturbed vegetation, and acquire blurred boundaries when images are resized. We developed AS-UNet, a lightweight U-Net variant with three targeted modifications. The Asymmetric Strip Attention Module uses horizontal and vertical depthwise strip convolutions with parallel channel-spatial reweighting to capture anisotropic landslide morphology. The Channel-Spatial Joint Gate uses decoder semantics to filter selected skip connections while retaining channel-specific spatial responses. The Poly-Harmonized Gradient Dice Loss (PGD Loss) combines pixel-wise, region-overlap, gradient-density, and probability-regularization terms for imbalanced segmentation. At 128 × 128 input resolution, AS-UNet achieved a best-validation IoU of 80.46 ± 0.03% and an independent-test IoU of 77.82 ± 0.32% across three random seeds. AS-UNet contains 8.634 M parameters and processed 380.79 frames per second on the reported hardware. These results indicate a favorable balance between segmentation accuracy and computational efficiency for RGB optical landslide mapping.
Haoran You, Cong Wang, Yu-di Qin et al.· Italian National Conference...· 0 citations
Object detection in optical remote sensing imagery is severely affected by adverse weather conditions, such as fog and haze, which degrade image quality and obscure structural details. Although recent Transformer-based detectors have achieved promising performance, they suffer from quadratic computational complexity for high-resolution inputs and tend to produce imprecise object boundaries in degraded scenes. To address these issues, we propose an Edge-Guided State-Space DETR (ES-DETR), an end-to-end detection framework that integrates linear-complexity state-space modeling with structural priors. Specifically, a Laplacian Edge-Aware Module (LEM) is designed to extract high-frequency boundary information from foggy images. Moreover, a Structural-Prior-Driven Mamba Fusion Module (SMF) is introduced to incorporate edge-derived structural priors into the Mamba architecture for long-range dependency modeling and feature fusion. This design effectively restores degraded semantic representations. Extensive experiments on foggy remote sensing benchmarks demonstrate that the proposed ES-DETR outperforms state-of-the-art detectors while maintaining a favorable accuracy–efficiency trade-off.
Xiaopeng Yang, Qiang Zhang, Zheng Liang et al.· IEEE Signal Processing Lette...· 0 citations
In karst regions, sugarcane mapping faces challenges from fragmented fields, undulating terrain, spectral confusion, and persistent cloud cover, which limit traditional optical remote sensing. To address these issues, we propose a fine-scale extraction framework that integrates Sentinel-2 optical and Sentinel-1 synthetic aperture radar (SAR) imagery through image-level fusion, and introduces a UNet-DFH network with a Multi-Scale Edge Fusion (MSEF) module and an Attention-Deformable Fusion Module (ADFM). This study makes three core contributions: (1) we construct a dedicated optical–SAR collaborative sugarcane extraction dataset for typical karst regions, alleviating the scarcity of multimodal labeled samples; (2) we propose the UNet-DFH network, where MSEF enhances boundary preservation and topological detail in shallow decoding stages, while ADFM improves robustness to geometric deformation and local misalignment in deep semantic stages; (3) we demonstrate that the joint mechanism of edge-preserving filtering and deformable adaptation yields a synergistic effect in addressing the precision–recall trade-off. Experiments in a typical karst area of Guangxi, China, demonstrate that optical–SAR fusion achieves an IoU of 80.08% and an OA of 92.09% during the sugar accumulation and maturity stage. During the more challenging tillering stage, UNet-DFH maintains relatively stable performance under optical-only conditions, with an IoU of 72.98%, Recall of 82.78%, and OA of 92.12%. Moreover, optical–SAR fusion improves Recall by 5.5 percentage points over optical-only inputs (from 83.54% to 89.04%), while Precision exhibits a moderate decrease from 91.89% to 88.84%, reflecting the expected trade-off associated with speckle noise. These results confirm the complementary value of multimodal data and the effectiveness of the proposed modules in preserving fragmented plot boundaries and improving segmentation performance in complex karst terrain. The framework offers a promising approach for high-precision crop mapping in the studied karst agricultural landscape.
Yanling Lu, Jinshuang Liu, Jingwen Li et al.· Remote Sensing· 0 citations
Multimodal semantic segmentation of high-resolution remote sensing imagery is important for fine-grained land-cover interpretation. However, existing fusion methods still suffer from unstable shallow optical-DSM alignment and deep feature degradation caused by heterogeneous frequency noise, boundary-detail loss, and inconsistent spatial responses. To address the aforementioned challenges, this letter proposes a coarse-to-fine progressive fusion network (CFPFNet). Specifically, a visual state space model extracts a Mamba-derived global structural prior to guide the coarse-grained context enhancement (CGCE) module for preliminary cross-modal alignment. Then, the fine-grained adaptive frequency-spatial fusion (FGAF) module performs amplitude-phase collaboration and adaptive spatial cross-gating for multiscale semantic refinement. Experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness of CFPFNet. On Vaihingen, CFPFNet improves mIoU and mF1 by 2.22% and 1.39% over the simple dual-stream baseline, respectively.
Di Zhang, Yuhang Yan, Q. Niu et al.· IEEE Geoscience and Remote S...· 2 citations
Timely and reliable mapping of landslide-affected areas from high-spatial-resolution optical imagery is essential for disaster investigation and post-event assessment. However, this task remains challenging because landslides usually exhibit large-scale variations, irregular boundaries, and strong spectral–textural similarities with surrounding bare-surface objects, which often cause missed detections, false positives, incomplete delineation, and inaccurate boundary localization. To address these problems, this paper presents a Scale-View Interactive Attention Network, named SIA-Net, for RGB-based landslide segmentation. First, a Multi-Scale Attention Module (MSAM) is constructed to encourage information exchange among features with different spatial resolutions. By doing so, the network can better represent both small scattered landslide patches and large continuous landslide bodies. Second, a Multi-View Attention Module (MVAM) is introduced to aggregate contextual cues from multiple receptive field views. This design strengthens the model’s ability to distinguish landslides from visually confusing objects, including bare soil, roads, riverbanks, and terrain shadows. In addition, a Convolutional Block Attention Module (CBAM) is incorporated during feature reconstruction to enhance landslide-related channel and spatial responses, thereby improving segmentation completeness and boundary localization. Experiments on the CAS Landslide Dataset (CLD) and GVLM Dataset show that SIA-Net provides more accurate landslide masks than the compared segmentation networks under the adopted benchmark settings. These results indicate that integrating scale-level interaction, view-level contextual modeling, and attention-guided decoding can effectively improve landslide extraction in complex optical remote sensing scenes.
High-resolution remote sensing semantic segmentation requires the joint modeling of local details, global semantics, and height-derived geometric structures, and it provides an important basis for urban object mapping, land-cover analysis, and fine-grained spatial understanding. However, in complex urban scenes, fine-grained boundaries, small objects, inter-class similarity, and spectral confusion can still weaken the stability of pixel-level prediction. To enhance discriminative dense feature representations in high-resolution remote sensing images, we propose HDSMNet, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data. The network separately extracts appearance and semantic features from optical imagery and height–structural features from nDSM, and introduces a Height-Guided Sparse Cross-Modal Fusion (HGSCF) module. Rather than treating nDSM as an additional feature source for generic fusion, HGSCF derives contextual representations, local feature contrasts, and structural-discontinuity cues from encoded nDSM features and uses them to guide sparse anchor-based interaction between optical and height features. This design enhances discriminative dense feature representations through interaction with a compact set of geometry-guided anchors. To complement HGSCF at the output stage, HDSMNet further adapts a Context-Guided Refinement (CGR) path that combines intermediate-response-guided contextual aggregation with dynamic feature modulation. This supplementary path recalibrates decoder features for output refinement. Experiments on the ISPRS Potsdam and Vaihingen datasets show that HDSMNet achieves mIoU values of 86.57% and 84.22%, respectively; ablation results further identify HGSCF as the main contributor to the observed improvement.
Unknown authors· Remote Sensing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.