Skip to content

Parameter-Efficient Local-Context Cooperation via Vision Foundation Models for UHR Remote Sensing Image Segmentation

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 4412214-4412214 · 0 citations · 66 references

Abstract

Ultrahigh resolution (UHR) remote sensing image segmentation aims to achieve a fine-grained understanding of complex ground scenes. In recent years, vision foundation models (VFMs) have shown strong capability in learning generic structural priors from large-scale visual data, indicating great potential for such fine-grained scene understanding. However, their application to UHR remote sensing images remains limited, as the massive parameter scales of VFMs are difficult to train under UHR remote sensing images. Motivated by the success of the parameter-efficient fine-tuning paradigm on VFMs, we propose a novel parameter-efficient local-context cooperation (PEACE) framework, which significantly reduces trainable parameter overhead while improving segmentation accuracy. In particular, PEACE leverages a shared VFM with minimal trainable parameters to collaboratively process local and corresponding contextual patches partitioned from the UHR remote sensing image. A multireceptive local adapter (MRLA) and a multireceptive context adapter (MRCA) are designed to capture spatial features of local and contextual inputs across multiple receptive fields. Finally, contextual semantics are integrated into local representations. Furthermore, a context-sensitive assistance strategy (CSAS) leverages the correct prediction of the context to effectively overcome the primary limitations of the patch-based training paradigm. Experimental results demonstrate that PEACE effectively exploits VFMs and remote sensing foundation models (RSFMs) with minimal parameter increments and achieves versatility and superior performance across several UHR remote sensing image benchmarks.

View source

Similar papers

2026

MsRE: Toward Efficient Remote Sensing Segmentation via Vision Foundation Models

Vision foundation models (VFMs) pretrained on large-scale datasets have significantly improved performance in remote sensing semantic segmentation. However, existing methods typically rely on full fine-tuning, which requires updating all model parameters. Instead of updating the full parameter set, parameter-efficient fine-tuning (PEFT) achieves competitive performance by optimizing only a small subset of parameters. Despite its success, most existing PEFT methods are mainly designed for natural image tasks and fail to account for the unique multiscale characteristics of remote sensing images. To address these challenges, we propose multi-scale cognitive feature refinement (MsRE) tuning, a novel PEFT method tailored for remote sensing semantic segmentation. In particular, MsRE captures multiscale contextual information by applying cognitive operations with different cognitive fields to intermediate features of the backbone. It then introduces a set of learnable tokens to establish interactions with features at different scales, enabling precise feature refinement and progressive feature propagation across network layers. This mechanism enhances the model’s ability to understand complex remote sensing scenes and improves downstream segmentation performance. With significantly fewer trainable parameters, MsRE provides an efficient yet effective solution for adapting VFMs to remote sensing segmentation tasks. Extensive experiments demonstrate that MsRE achieves competitive segmentation performance with substantially fewer trainable backbone parameters, providing a favorable balance between accuracy and parameter efficiency. The project is available at http://woldier.top/MsRE

Bin Wang, Shun Lv, Zhi Li et al. · 0 citations
2026

A Coarse-to-Fine Progressive Fusion Network for Multimodal Semantic Segmentation of Remote Sensing Images

Multimodal semantic segmentation of high-resolution remote sensing imagery is important for fine-grained land-cover interpretation. However, existing fusion methods still suffer from unstable shallow optical-DSM alignment and deep feature degradation caused by heterogeneous frequency noise, boundary-detail loss, and inconsistent spatial responses. To address the aforementioned challenges, this letter proposes a coarse-to-fine progressive fusion network (CFPFNet). Specifically, a visual state space model extracts a Mamba-derived global structural prior to guide the coarse-grained context enhancement (CGCE) module for preliminary cross-modal alignment. Then, the fine-grained adaptive frequency-spatial fusion (FGAF) module performs amplitude-phase collaboration and adaptive spatial cross-gating for multiscale semantic refinement. Experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness of CFPFNet. On Vaihingen, CFPFNet improves mIoU and mF1 by 2.22% and 1.39% over the simple dual-stream baseline, respectively.

Di Zhang, Yuhang Yan, Q. Niu et al. · 2 citations
Open access Aug 2026

Towards Lightweight and Accurate Remote-Sensing Image Super-Resolution via Reparameterized Feature Enhancement Network

Remote sensing image super-resolution (RSISR) provides an effective means of improving spatial detail for Earth observation and satellite image interpretation. However, existing methods often rely on increasingly complex network designs with deeper hierarchies and expanded channel capacities to pursue higher performance, resulting in heavy models with high computational cost, which restricts their deployment on resource-constrained platforms. To address this challenge, we propose a novel reparameterized feature enhancement network (RepFEN) for lightweight and accurate RSISR tasks. Specifically, a multi-scale reparameterized module (MRepM) is designed to capture multi-scale spatial information and enhance texture representation. Furthermore, a partial-channel gated attention module (PCGAM) is introduced to selectively enhance discriminative features along the channel dimension, effectively improving fine-grained detail restoration. By integrating structural reparameterization and multi-scale lightweight modules, the proposed method achieves a better balance between reconstruction accuracy and inference efficiency. Extensive experiments on both remote sensing and natural image super-resolution benchmarks demonstrate that our method achieves superior performance compared to existing state-of-the-art methods, while maintaining minimal computational overhead, showing significant potential for real-world applications.

Feng Huang, Ren-Hui Wei, Liqiong Chen et al. · 0 citations
Preprint Aug 2026

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.

Shuai Cao, Meng Tang, Shuwei Peng et al. · 0 citations
2026

CFMNet: A Lightweight Backbone With Cooperative Feature Modeling for Remote Sensing Vision Tasks

High-resolution remote sensing imagery presents unique challenges for efficient visual understanding, including dense object distributions, severe scale variations, strong background redundancy, and complex spatial structures. Existing deep models often rely on deep architectures or computationally intensive global modeling strategies, limiting their deployment on resource-constrained platforms. In this article, we propose an efficient and lightweight backbone network, termed the cooperative feature modeling network (CFMNet), for high-resolution remote sensing image understanding. CFMNet decomposes feature representations into heterogeneous yet complementary subspaces and models them cooperatively within a unified framework. Specifically, it coordinates channel semantics, structure-aware spatial dependencies, local detail enhancement, and global contextual consistency to improve representation efficiency while suppressing redundant computation. Extensive experiments demonstrate the effectiveness and generality of CFMNet. It achieves 96.13%, 95.50%, and 98.10% Top-1 accuracy on NWPU-RESISC45, aerial image dataset (AID), and UC Merced Land Use dataset (UCM), respectively, 79.82% mAP on DOTA-v1.0, 73.02% mAP on DOTA-v1.5, and 90.82% mAP on HRSC2016, as well as 83.8% mIoU on Vaihingen and 53.8% mIoU on LoveDA, while maintaining low parameter count and computational complexity. A scaling-based Pareto analysis on DOTA-v1.0 and LoveDA further shows that CFMNet variants form a favorable efficiency–accuracy frontier compared with representative lightweight backbones. These results indicate that cooperative modeling of heterogeneous features provides an effective and efficient solution for high-resolution remote sensing image understanding. The code will be released at https://github.com/BEIBEIPRINCESS/CFMNet

Jih-Ming Chen, Haonan Guo, Jun Liu et al. · 0 citations
Open access Aug 2026

Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation

Frequency and Edge-guided SAM (FE-SAM) is proposed, a scalable and efficient framework for RSISS that adaptively decomposes and modulates frequency-domain features based on the input data and designs EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image.

Feng Gao, Zizhe Pan, Haoting Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.