Skip to content
Review Open access

Foundation Models for Remote Sensing Semantic Segmentation: A Review of Architectures, Adaptations, and Prospects

Jul 2026 · Italian National Conference on Sensors · Vol 26, pp. 4671 · 0 citations · 124 references
Medicine

TL;DR

This paper reviews recent progress from the perspectives of dataset evolution, model architectures, and downstream adaptation strategies, covering parameter-efficient fine-tuning, prompt engineering, few-shot and zero-shot learning, open-vocabulary segmentation, and domain adaptation.

Abstract

Highlights What are the main findings? Review of four emerging foundation model paradigms for remote sensing image segmentation—Transformer-based architectures, state space models (Mamba), prompt-driven segmentation (SAM), and self-supervised or multimodal pre-training analyzing trade-offs in global context modeling, computational efficiency, and cross-modal representation. Synthesis of downstream adaptation strategies, including parameter-efficient fine-tuning (LoRA, adapters), prompt engineering, few-shot and zero-shot learning, open-vocabulary segmentation, and domain adaptation, revealing how each strategy addresses the gap between pre-training and remote sensing requirements. What are the implications of the main findings? Identification of fundamental bottlenecks limiting current models, including the tension between representation generality and remote sensing-specific adaptation, multimodal sensor heterogeneity, and insufficiencies in existing evaluation ecosystems and annotation paradigms. A forward-looking research roadmap toward remote-sensing-native pre-training, lightweight edge-deployable architectures, and unified open-world geospatial foundation models, providing guidance for future research and practical deployment. Abstract Remote sensing image segmentation is a foundational task in Earth observation. With the rapid growth of remote sensing datasets in terms of scale, modality diversity, semantic openness, and spatio-temporal complexity, the field is evolving from task-specific supervised learning toward foundation-model paradigms. Recent advances in foundation models—including Transformer-based architectures, Mamba-based state space models (SSMs), prompt-driven frameworks such as the Segment Anything Model (SAM), and self-supervised or multimodal pre-training—have profoundly reshaped the technical landscape of remote sensing image segmentation. This paper reviews recent progress from the perspectives of dataset evolution, model architectures, and downstream adaptation strategies, covering parameter-efficient fine-tuning, prompt engineering, few-shot and zero-shot learning, open-vocabulary segmentation, and domain adaptation. We further analyze core challenges including the tension between representation generality and remote sensing-specific adaptation, multimodal sensor heterogeneity, and the insufficiency of existing evaluation ecosystems. Finally, we discuss future directions toward remote-sensing-native pre-training, lightweight edge deployment, and unified open-world geospatial foundation models.

Read PDF

Similar papers

2026

MsRE: Toward Efficient Remote Sensing Segmentation via Vision Foundation Models

Vision foundation models (VFMs) pretrained on large-scale datasets have significantly improved performance in remote sensing semantic segmentation. However, existing methods typically rely on full fine-tuning, which requires updating all model parameters. Instead of updating the full parameter set, parameter-efficient fine-tuning (PEFT) achieves competitive performance by optimizing only a small subset of parameters. Despite its success, most existing PEFT methods are mainly designed for natural image tasks and fail to account for the unique multiscale characteristics of remote sensing images. To address these challenges, we propose multi-scale cognitive feature refinement (MsRE) tuning, a novel PEFT method tailored for remote sensing semantic segmentation. In particular, MsRE captures multiscale contextual information by applying cognitive operations with different cognitive fields to intermediate features of the backbone. It then introduces a set of learnable tokens to establish interactions with features at different scales, enabling precise feature refinement and progressive feature propagation across network layers. This mechanism enhances the model’s ability to understand complex remote sensing scenes and improves downstream segmentation performance. With significantly fewer trainable parameters, MsRE provides an efficient yet effective solution for adapting VFMs to remote sensing segmentation tasks. Extensive experiments demonstrate that MsRE achieves competitive segmentation performance with substantially fewer trainable backbone parameters, providing a favorable balance between accuracy and parameter efficiency. The project is available at http://woldier.top/MsRE

Bin Wang, Shun Lv, Zhi Li et al. · 0 citations
Open access Jul 2026

Leveraging Pretrained Priors for Weakly Supervised Semantic Segmentation of Remote Sensing Images

A lightweight and efficient framework that integrates CLIP and DINO foundation models to address three challenges: semantic misalignment between generic text prompts and RSI-specific visuals; static CAM quality; and incomplete object coverage is proposed.

Xin Li, Nicola Genzano, M. Gianinetto et al. · 0 citations
Preprint Aug 2026

BASeg: Boundary-Aware Remote Sensing Segmentation with Structural Penalties

A Mahalanobis-Angle Boundary Loss (MABL) is proposed that explicitly enhances boundary and shape consistency and is introduced, built upon MABL, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties.

Yuexi Song, Kailai Sun, Zhuoyue Wang et al. · 0 citations
Sep 2026

PromptRefine: source-free unsupervised domain adaptation for remote sensing image classification via CLIP filtering

Despite the success of deep learning in remote sensing (RS) image classification, substantial domain shifts—stemming from heterogeneous sensors and diverse environmental conditions—frequently compromise model reliability. Although source-free unsupervised domain adaptation (SFUDA) has emerged as a critical paradigm to bypass data privacy and storage constraints, existing methods remain fragile in complex RS scenes where noisy pseudo-labels often trigger catastrophic semantic drift. We propose PromptRefine, a white-box SFUDA framework designed to anchor target adaptation through cross-modal intelligence. Specifically, we leverage the zero-shot semantic priors of large-scale vision–language models (e.g., contrastive language–image pre-training) to rectify source-biased predictions via a dynamic prompt fine-tuning mechanism. The framework executes a three-stage alternating optimization strategy that integrates cross-modal semantic alignment, hard-sample mining via sliced Wasserstein discrepancy, and fine-grained prompt evolution. Evaluated across 18 cross-domain tasks on five benchmarks (UCM, WHU-RS19, AID, RSSCN7, and NWPU-RESISC45), PromptRefine consistently achieves remarkable performance. Specifically, it outperforms the leading vision transformer-based baseline (VisTA) by 0.42% and 0.47% on the two groups of cross-domain tasks and surpasses the top ResNet-based baseline (SRKT/DFENet) by 2.94% and 5.02% in average classification accuracy. The proposed method also outperforms other SFUDA and unsupervised domain adaptation methods on all 18 tasks, demonstrating its superior adaptation capability. Our approach provides a robust, privacy-preserving, and computationally efficient solution, setting a benchmark for scalable RS scene characterization.

Unknown authors · 0 citations
2026

Text-Guided Dual Refinement for Domain Generalized Semantic Segmentation in Remote Sensing

Recently, domain generalized remote sensing semantic segmentation (DG-RSSS) methods leverage vision foundation models (VFMs) with a parameter-efficient fine-tuning (PEFT) strategy to achieve remarkable progress. Although VFMs offer robust representations under distribution shifts between different remote sensing scenes, they still have some limitations, including inaccurate segmentation between similar classes and boundary regions. To address these limitations, this work proposes a text-guided dual refinement (TGDR) approach for DG-RSSS, which contains a text-guided discriminative refinement (TDR) module and a text-guided mask features refinement (TMR) module. In particular, first, the proposed TDR module generates class-discriminated features by using class-related learnable tokens and interclass similarity to refine the original features from the frozen backbone, where the tokens are initialized with class texts. Second, the proposed TMR module injects class-specific semantics into the phase component of mask features by using class-related learnable tokens to enhance the correlation between the semantics of classes and scene contents in mask features for refining the prediction of class boundaries in scene contents. Extensive experiments demonstrate that the proposed TGDR approach achieves superior performance across multiple DG-RSSS benchmarks, e.g., achieving 65.6%, 54.0%, 59.0%, 45.2%, 48.8%, and 79.3% mIoU on the Potsdam–to–Vaihingen (P2V), Vaihingen–to–Potsdam (V2P), Rural–to–Urban (R2U), Urban–to–Rural (U2R), Aerial–to–Satellite II (A2S), and Satellite II–to–Aerial (S2A) benchmarks.

Muxin Liao, Mei-Ying Liao, Yuting Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.