2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 4413222-4413222· 0 citations· 69 references
Abstract
Semantic segmentation of remote sensing images (RSIs) often struggles with boundary blurring, structural discontinuity, and category confusion due to limitations in conventional interpolation and dynamic upsampling methods. This article proposes RSUS, a new upsampling layer designed to preserve semantic consistency and spatial structure. RSUS consists of three components: 1) global context vector aggregation (GCVA) for content-aware prediction kernels that introduce global priors into local feature reconstruction; 2) cross-scale anisotropic implicit positional encoding (CAIPE) for direction-sensitive spatial deformation and structural alignment; and 3) adaptive high-frequency structure gating (AHSG) to enhance boundary-related frequency responses. In addition, persistent homology (PH) is used as a topological analysis tool to assess connectivity and structure preservation beyond pixel-level metrics. Extensive experiments on five datasets (International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen, ISPRS Potsdam, LoveDA, UAVid, and Xining) demonstrate that RSUS improves performance across mainstream segmentation networks, outperforming advanced upsampling methods. Analyses show RSUS alleviates structural misalignment, improves category consistency, and recovers fine-grained boundaries in remote sensing segmentation. The source code is available at: https://github.com/Ronin-711/RSUS
The semantic interpretation of remote sensing imagery through segmentation has become indispensable for a wide range of applications, including resource exploration, environmental assessment, and land-use analysis. Yet, accurate parsing of such images remains challenging because complex object boundaries and large scale differences often weaken the ability of conventional Convolutional Neural Network (CNN)-based methods to preserve local details. In response, this study constructs a segmentation framework that couples wavelet convolution with the Mamba architecture. To strengthen feature learning in the intermediate stages, an Auxiliary Segmentation Module (ASM) is employed to provide additional supervisory guidance, which supports optimization and encourages the representation of subtle semantic details. Wavelet-transform convolution is also introduced into the downsampling path, enabling spatial cues and frequency-related information to be exploited in a more coordinated manner for finer boundary and texture modeling. Experiments on public remote sensing datasets and mining area imagery further confirm the effectiveness of the method. Compared with several existing segmentation approaches, the proposed model delivers better overall performance in mIoU, F1-score, and recognition accuracy, particularly in scenes where multiple land-cover categories are heavily interlaced. Moreover, these gains are obtained with relatively low model complexity, suggesting good potential for practical deployment in land monitoring and ecological management.
Wenxi He, Zongmin Yin, Yulong Yang et al.· Remote Sensing· 0 citations
Frequency and Edge-guided SAM (FE-SAM) is proposed, a scalable and efficient framework for RSISS that adaptively decomposes and modulates frequency-domain features based on the input data and designs EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image.
Feng Gao, Zizhe Pan, Haoting Wang et al.· IEEE Transactions on Geoscie...· 0 citations
A Mahalanobis-Angle Boundary Loss (MABL) is proposed that explicitly enhances boundary and shape consistency and is introduced, built upon MABL, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties.
Yuexi Song, Kailai Sun, Zhuoyue Wang et al.· 0 citations
Road-network extraction from very high-resolution (VHR) remote-sensing imagery remains a challenging task owing to the structural sparsity, topological complexity, and severe occlusions of road networks. Conventional graph-based approaches preserve topological consistency yet incur considerable computational overhead, whereas prevailing convolutional neural network (CNN) and Transformer architectures struggle to reconcile long-range contextual modeling with computational efficiency. To address these limitations, this study proposes RFM-UNet, a hybrid frequency and state–space network designed for road-network segmentation. Specifically, the encoder integrates Mamba blocks with an Anisotropic Directional Attention (ADA) module to jointly capture local geometric cues and global dependencies at linear computational complexity. In addition, a Multi-Scale Adaptive Fusion Module (MAFM) is introduced to dynamically recalibrate multi-stage features, thereby suppressing cross-scale interference and preserving the connectivity of narrow roads. To enhance robustness against shadow-induced occlusions, a Dual-Spectrum Aggregation Module (DualSpec) decouples the phase and amplitude spectra in the frequency domain and fuses them with spatial features, effectively mitigating spurious responses and background noise characterized by similar textures. Quantitative and qualitative experiments on three public datasets demonstrate that RFM-UNet consistently outperforms current state-of-the-art methods.
Pu Song, Peng Yu, Xiaojing Zhong et al.· Remote Sensing· 0 citations
Semantic segmentation of remote sensing imagery has been widely applied in landslide identification, effectively addressing the time-consuming and labor-intensive nature of manual visual interpretation. However, existing models still face challenges in extracting multiscale features and accurately delineating boundaries under complex background conditions. To overcome these limitations, this study proposes a multiscale boundary aware network (MBANet) for landslide identification in remote sensing imagery. Specifically, we design a multiscale cross-interaction convolution (MCC) module that captures local details and broader contextual cues through heterogeneous receptive-field branches, and recalibrates the concatenated multiscale features via an adaptive cross-branch interaction strategy. In addition, a boundary sensitive refinement attention (BSRA) module is introduced to enhance boundary localization by combining a Sobel-based gradient prior, learnable boundary estimation, and region-context enhancement for fine-grained boundary refinement. These modules are integrated into an encoder–decoder architecture to jointly achieve semantic consistency and boundary precision. Experimental results on two public datasets show that MBANet outperforms other comparison models in overall segmentation performance and maintains competitive performance in boundary delineation. On the Bijie dataset, it achieves a recall of 83.84% and an $F1$ -score of 85.39%; on the Palu dataset, it reaches a recall of 75.83% and an $F1$ -score of 78.17%, highlighting its superior performance.
Zixun Xie, Chuang Song, Xingmin Cai et al.· IEEE Geoscience and Remote S...· 0 citations
DeepKANSeg, a novel network based on the Kolmogorov–Arnold network (KAN), achieves superior performance in terms of accuracy compared to state-of-the-art methods and improves interpretability, making it suitable for explainable learning in remote sensing.
Ziyao Wang, Yin Hu, Xiaokang Zhang et al.· IEEE Transactions on Geoscie...· 12 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.