Skip to content

S2CLNet: Structure-Constrained Semantic Contrastive Learning for Referring Remote Sensing Image Segmentation

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5636114-5636114 · 0 citations · 60 references

Abstract

Referring remote sensing image segmentation (RRSIS) aims to segment target ground objects in high-resolution remote sensing imagery according to textual descriptions. Due to complex backgrounds and large-scale variations, accurately aligning linguistic semantics with spatial regions remains a challenging problem. Existing RRSIS methods are predominantly segmentation-centric, where textual descriptions mainly serve as conditional guidance for pixel-wise prediction rather than explicitly enforcing semantic consistency between visual regions and linguistic representations. This limitation often leads to suboptimal fine-grained alignment and insufficient suppression of background-dominant regions. To address these challenges, we propose a novel RRSIS framework centered on structure-constrained semantic contrastive learning (CL), termed S2CLNet. Departing from the segmentation-driven paradigm, instead of relying solely on segmentation supervision, S2CL introduces region-level structural constraints to perform fine-grained CL in the joint embedding space. This design encourages semantically corresponding visual regions and linguistic representations to be more closely aligned while effectively separating target regions from irrelevant background areas. Furthermore, we develop a dynamic modality balancing module (DMBM) to enhance cross-modal interaction under complex remote sensing scenarios. The DMBM jointly models bidirectional cross-modal attention (BCA) and dynamically adjusts the relative contributions of visual and linguistic modalities according to scene complexity and linguistic specificity, thereby facilitating more adaptive and robust multimodal feature integration. Extensive experiments on three public benchmarks, RefSegRS, RRSIS-D, and RISBench, demonstrate that the proposed method achieves superior performance compared with state-of-the-art approaches. The code will be publicly available at https://github.com/45degreesl

View source

Similar papers

Open access Jul 2026

Leveraging Pretrained Priors for Weakly Supervised Semantic Segmentation of Remote Sensing Images

A lightweight and efficient framework that integrates CLIP and DINO foundation models to address three challenges: semantic misalignment between generic text prompts and RSI-specific visuals; static CAM quality; and incomplete object coverage is proposed.

Xin Li, Nicola Genzano, M. Gianinetto et al. · 0 citations
Open access Aug 2026

Weakly Supervised Remote Sensing Segmentation via Decoupled Cross-Modal Distillation and Semantic-Guided Refinement

A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing scenarios under image-level supervision.

Jing Li, Yulin Cao, Xiantao Jiang et al. · 0 citations
Preprint Aug 2026

Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing

DinoSplat-OV is proposed, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.

Changhao Zhao, Haoxiang Li, Yuke Li et al. · 0 citations
Open access 2026

SAPLNet: State-Aware Prototype Learning for Remote Sensing Segmentation

Semantic segmentation of high-resolution remote sensing images remains challenging due to complex spatial structures, multiscale object variations, fine-grained category differences, and high interclass similarities. Conventional segmentation methods usually rely on fixed convolutional heads or single feature representations, which makes it difficult to effectively model both intraclass appearance variations and interclass texture similarities, often leading to category confusion, missed objects, and incomplete segmentation in complex scenes. To address these challenges, we propose a state-aware prototype learning network, termed SAPLNet. Specifically, a cross-stage state refiner is introduced to progressively refine multilevel features by integrating the input features with the outputs of different stages through state-aware gated normalization. Then, a weighted feature pyramid decoder performs top-down fusion of the refined hierarchical features, combining high-level semantic information with low-level spatial details. Furthermore, a state-aware multiprototype classifier is designed to construct multiple semantic prototypes for each class via ground-truth-guided local class-center extraction and momentum-based prototype memory updating. A global state vector derived from the refined cross-stage features is used to adaptively modulate decoder features, improving the matching reliability between pixel features and class prototypes. In addition, prototype compactness loss, prototype diversity loss, and lightweight boundary loss are employed to enhance intraclass consistency, prototype discriminability, and boundary awareness. Experimental results demonstrate the effectiveness and superiority of SAPLNet.

Zeyu Zhao, Zhaolong Gao, Jun Feng · 0 citations
2026

Graph2Scene: Generating Remote Sensing Imagery and Labels via Scene Graphs and Low-Rank Representation

High-precision, pixel-level annotations are indispensable for remote sensing semantic segmentation and related tasks, yet producing such labels manually is prohibitively expensive. Although recent generative models can synthesize realistic remote sensing data, existing approaches typically either rely heavily on preexisting ground-truth masks as conditioning inputs or lack precise control over the spatial layout of the generated content. To address this gap, we propose Graph2Scene, a novel framework for the joint generation of remote sensing images and pixel-level labels driven by scene graphs. This framework establishes a flexible control mechanism that utilizes scene graphs derived from existing semantic labels during training to learn semantic priors, while enabling users to explicitly define object quantities, categories, and topological relationships for customized generation during inference. Graph2Scene adopts a two-stage cascade: Graph2Mask encodes the scene graph into textual prompts and employs an image-level low-rank adaptation (LoRA) to finetune a FLUX model for label generation; Mask2Scene uses a class-level LoRA strategy to learn fine-grained visual features and generates remote sensing images conditioned on the label. Experiments on our developed non-agricultural conversion process (NACP) dataset and the public LoveDA dataset show that Graph2Scene effectively produces structurally coherent and visually realistic remote sensing images together with accurate pixel-level annotations. The code will be made available at https://github.com/GeoRSAI/Graph2Scene

Shaoxuan Zhao, Xiaoguang Zhou, Dongyang Hou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.