Aug 2026· IEEE Transactions on Geoscience and Remote Sensing· 0 citations· 55 references
Computer Science
TL;DR
Frequency and Edge-guided SAM (FE-SAM) is proposed, a scalable and efficient framework for RSISS that adaptively decomposes and modulates frequency-domain features based on the input data and designs EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image.
Abstract
Remote sensing image semantic segmentation (RSISS) has attracted significant attention due to the growing demand for fine-grained land cover information. The Segment Anything Model (SAM), proposed as a foundation vision model, offers strong segmentation performance and generalization capabilities for RSISS tasks. However, existing SAM-based approaches face two limitations: (1) Insufficient adaptation of SAM's features to the diverse characteristics of land cover types. (2) Semantic ambiguity at object boundaries, which hinders accurate delineation. To address these limitations, we propose Frequency and Edge-guided SAM (FE-SAM), a scalable and efficient framework for RSISS. Specifically, we introduce a Frequency-Modulated Adapter (FMA) that adaptively decomposes and modulates frequency-domain features based on the input data. It selectively enhances informative high- and low-frequency components corresponding to different land cover types. Furthermore, to improve SAM's ability to capture fine-grained details, we design EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image. Extensive experiments on three benchmark datasets demonstrate that FE-SAM outperforms state-of-the-art methods. The source codes are available at: https://github.com/oucailab/FE-SAM.
A Mahalanobis-Angle Boundary Loss (MABL) is proposed that explicitly enhances boundary and shape consistency and is introduced, built upon MABL, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties.
Yuexi Song, Kailai Sun, Zhuoyue Wang et al.· 0 citations
The results demonstrate that AB-SAM provides a practical parameter-efficient framework for automated, hint-free landslide segmentation, although further evaluation across additional regions, sensors, and landslide-size distributions remains necessary.
This work presents a new method for using the VFM Segment Anything Model 2 (SAM 2) for multi-class semantic segmentation of Sentinel-2 images that does not require training data and achieves an overall accuracy of up to 93% at pixel-level using polygon mask prompts.
Paula L. Lippmann, M. Dorozynski, F. Rottensteiner et al.· The International Archives o...· 0 citations
Vision foundation models (VFMs) pretrained on large-scale datasets have significantly improved performance in remote sensing semantic segmentation. However, existing methods typically rely on full fine-tuning, which requires updating all model parameters. Instead of updating the full parameter set, parameter-efficient fine-tuning (PEFT) achieves competitive performance by optimizing only a small subset of parameters. Despite its success, most existing PEFT methods are mainly designed for natural image tasks and fail to account for the unique multiscale characteristics of remote sensing images. To address these challenges, we propose multi-scale cognitive feature refinement (MsRE) tuning, a novel PEFT method tailored for remote sensing semantic segmentation. In particular, MsRE captures multiscale contextual information by applying cognitive operations with different cognitive fields to intermediate features of the backbone. It then introduces a set of learnable tokens to establish interactions with features at different scales, enabling precise feature refinement and progressive feature propagation across network layers. This mechanism enhances the model’s ability to understand complex remote sensing scenes and improves downstream segmentation performance. With significantly fewer trainable parameters, MsRE provides an efficient yet effective solution for adapting VFMs to remote sensing segmentation tasks. Extensive experiments demonstrate that MsRE achieves competitive segmentation performance with substantially fewer trainable backbone parameters, providing a favorable balance between accuracy and parameter efficiency. The project is available at http://woldier.top/MsRE
Bin Wang, Shun Lv, Zhi Li et al.· IEEE Transactions on Geoscie...· 0 citations
Semantic segmentation of remote sensing images (RSIs) often struggles with boundary blurring, structural discontinuity, and category confusion due to limitations in conventional interpolation and dynamic upsampling methods. This article proposes RSUS, a new upsampling layer designed to preserve semantic consistency and spatial structure. RSUS consists of three components: 1) global context vector aggregation (GCVA) for content-aware prediction kernels that introduce global priors into local feature reconstruction; 2) cross-scale anisotropic implicit positional encoding (CAIPE) for direction-sensitive spatial deformation and structural alignment; and 3) adaptive high-frequency structure gating (AHSG) to enhance boundary-related frequency responses. In addition, persistent homology (PH) is used as a topological analysis tool to assess connectivity and structure preservation beyond pixel-level metrics. Extensive experiments on five datasets (International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen, ISPRS Potsdam, LoveDA, UAVid, and Xining) demonstrate that RSUS improves performance across mainstream segmentation networks, outperforming advanced upsampling methods. Analyses show RSUS alleviates structural misalignment, improves category consistency, and recovers fine-grained boundaries in remote sensing segmentation. The source code is available at: https://github.com/Ronin-711/RSUS
Yaning Liu, Ronghao Yang, Shaoda Li et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.