Skip to content
Preprint

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

The ISRS-DETR employs an RF-DETR decoder with the interactive segmentation backbone to localise co-occurring same-class objects, and introduces a Dynamic Top-K Click Selection strategy that retains only reliable proposals and converts each into a simulated click, so one user interaction propagates across an entire class.

Abstract

Interactive segmentation reduces the prohibitive cost of pixel-level annotation by allowing users to delineate objects with a few clicks. However, applying this paradigm directly to remote sensing imagery is non-trivial: ultra-high resolutions, small object sizes, and sparse spatial distributions all degrade segmentation quality. Recent work has addressed the resolution barrier and achieved competitive results in interactive segmentation for remote sensing (ISRS). However, they treat all instances of a class within an image as a single objective target. Consequently, interactions spent on one object contribute nothing to its same-class neighbours, and satisfactory masks may demand up to 40 clicks per image, hindering the practicality of these frameworks. We observe that remote sensing scenes exhibit markedly strong inter-object correlation, meaning a single clicked object is highly informative about the rest of its category. Building on this, we propose ISRS-DETR, a detection-guided interactive segmentation framework that injects object-level evidence into both training and inference. Our ISRS-DETR employs an RF-DETR decoder with the interactive segmentation backbone to localise co-occurring same-class objects, and introduces a Dynamic Top-K Click Selection strategy that retains only reliable proposals and converts each into a simulated click, so one user interaction propagates across an entire class. Experiments on three standard remote sensing benchmarks show that ISRS-DETR achieves state-of-the-art accuracy while substantially reducing Number of Clicks per Image (NoC-I). All codes and data splits will be released for reproducibility upon acceptance.

View source

Similar papers

Open access Aug 2026

Weakly Supervised Remote Sensing Segmentation via Decoupled Cross-Modal Distillation and Semantic-Guided Refinement

A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing scenarios under image-level supervision.

Jing Li, Yulin Cao, Xiantao Jiang et al. · 0 citations
Open access Aug 2026

Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation

Frequency and Edge-guided SAM (FE-SAM) is proposed, a scalable and efficient framework for RSISS that adaptively decomposes and modulates frequency-domain features based on the input data and designs EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image.

Feng Gao, Zizhe Pan, Haoting Wang et al. · 0 citations
2026

CGSNet: Category Prior-Guided Self-Supervised Semantic Segmentation for Remote Sensing Images

Multimodal fusion methods have shown great potential in remote sensing image analysis, but existing approaches rely heavily on massive amounts of annotated data. This is not only costly and time-consuming but also prone to subjective bias. To address this issue, we propose a category-prior-based self-supervised framework, CGSNet, which uses category prior maps extracted from multispectral images as supervisory signals for end-to-end training. An adaptive confidence-weighted pseudo-label generation mechanism is designed to alleviate noise and errors in prior maps by replacing binary labels with continuous confidence maps, enabling the learning of uncertain interclass features. In addition, a multispectral feature-guided refinement strategy utilizes color and texture information to calibrate class transition regions and enhance the discriminative power of pseudo-labels in complex scenes. A dynamic mask selection strategy further enhances the model’s robustness and generalization capabilities through progressive learning. Experiments demonstrate that CGSNet achieves state-of-the-art performance without the need for human annotation, achieving an Mean Intersection over Union (mIoU) score of 78.46% on the Gaofen image dataset (GID) (vegetation) dataset and 79.58% on the Zurich (vegetation) dataset—12.22% and 15.07% higher than existing methods, respectively—while exhibiting strong cross-dataset zero-shot generalization capabilities. The code will be available at https://github.com/NUAALISILab

Jiahang Liu, Jian Cui, Mao-yin Guo et al. · 0 citations
Open access 2026

SAPLNet: State-Aware Prototype Learning for Remote Sensing Segmentation

Semantic segmentation of high-resolution remote sensing images remains challenging due to complex spatial structures, multiscale object variations, fine-grained category differences, and high interclass similarities. Conventional segmentation methods usually rely on fixed convolutional heads or single feature representations, which makes it difficult to effectively model both intraclass appearance variations and interclass texture similarities, often leading to category confusion, missed objects, and incomplete segmentation in complex scenes. To address these challenges, we propose a state-aware prototype learning network, termed SAPLNet. Specifically, a cross-stage state refiner is introduced to progressively refine multilevel features by integrating the input features with the outputs of different stages through state-aware gated normalization. Then, a weighted feature pyramid decoder performs top-down fusion of the refined hierarchical features, combining high-level semantic information with low-level spatial details. Furthermore, a state-aware multiprototype classifier is designed to construct multiple semantic prototypes for each class via ground-truth-guided local class-center extraction and momentum-based prototype memory updating. A global state vector derived from the refined cross-stage features is used to adaptively modulate decoder features, improving the matching reliability between pixel features and class prototypes. In addition, prototype compactness loss, prototype diversity loss, and lightweight boundary loss are employed to enhance intraclass consistency, prototype discriminability, and boundary awareness. Experimental results demonstrate the effectiveness and superiority of SAPLNet.

Zeyu Zhao, Zhaolong Gao, Jun Feng · 0 citations
Sep 2026

StructPointNet: explicit geometric prior extraction and reuse for efficient multimodal remote sensing semantic segmentation

Semantic segmentation is widely regarded as one of the most advanced scene understanding techniques in remote sensing. Although monomodal segmentation models have advanced recently, the rich potential of multimodal data is still largely untapped. Current multimodal approaches are often rigid, typically limited to bimodal inputs, and failing to adapt flexibly when three or more co-registered two-dimensional grid-based modalities are available. A more pressing issue is the computational burden inherent in traditional fusion paradigms; as the number of input modalities increases, parameter counts and processing loads often increase substantially, making them less practical for scalable multi-source remote sensing interpretation. Thus, creating a scalable and lightweight framework for co-registered grid-based multimodal remote sensing data remains an open challenge. To bridge this gap, we introduce StructPointNet, a lightweight segmentation framework built on the philosophy of “explicit extraction and reuse of geometric priors.” Our approach balances efficiency with precision through three core mechanisms: the structure-sensitive modality encoder, which captures modality-specific high-frequency geometric details via parallel Sobel branches and edge-guided attention; the Heterogeneity rectification layer and pyramidal hybrid backbone, which map diverse features into a shared latent space for global-local context modeling; and the boundary-aware point decoder, which refines boundary segmentation by resampling shallow structural features based on uncertainty estimates. Thanks to these designs, StructPointNet functions as a flexible framework for monomodal and multimodal segmentation with variable numbers of co-registered grid-based inputs while keeping the additional cost of each lightweight modality-specific branch controlled compared with conventional multistream backbone replication. We benchmarked our approach against multiple representative state-of-the-art models from the last five years on two public multimodal datasets. Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, particularly in defining complex geo-object boundaries. Moreover, the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.

Siwei Wei, Ruoxi Wang, Xueli Chang · 0 citations
Preprint Aug 2026

Pixel Ignores, Superpixel Sees: Adverse Weather Image Restoration via Semantic-Center SSM

Adverse weather image restoration aims to recover clear visibility from degraded images in complex weather conditions. Existing works attempt to address this problem by modeling relationships between pixels, however, this paradigm defies the spatially non-uniformity fact of degradations and learns non-discriminative features from semantic-conflict regions. In this paper, we propose SSR, a \textbf{S}emantic-center guilded \textbf{S}tate space model for image \textbf{R}estoration. The key idea of SSR is to shift the conventional scanning strategy of pixel-serial to semantic-guilded one. Specifically, we introduce a Superpixel-guided Selective Scan Mechanism ($\text{S}^3$M), which first partitions the image into perceptually coherent regions via superpixel clustering and then performs relations modeling within the semantic-related regions. Moreover, a Region-level Gating Mechanism (RGM) is developed to perform intra-region calibration by modulating degradation outliers within each semantic superpixel unit along the channel dimension. Extensive experiments on \textbf{6} well-established benchmarks demonstrate that SSR performs favorably against state-of-the-art models with competitive computational cost.

Dayu Li, Shihao Zhou, Leizhi Shu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.