Skip to content
Preprint

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

GeoSeg-OV is proposed, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding, and introduces Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context.

Abstract

Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization. Recent attempts have begun to incorporate auxiliary vision foundation models (VFMs), typically coupling their features with text embeddings as additional matching evidence. However, this strategy may introduce inconsistent matching signals while leaving the structure-sensitive representations of VFMs insufficiently exploited. We therefore propose GeoSeg-OV, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding. GeoSeg-OV constructs an orientation-robust cost volume from multi-rotation CLIP features, while a frozen VFM extracts multi-scale structure-sensitive features in parallel. We propose Structure-Guided Aggregation (SGA), which integrates cost tokens and CLIP semantic guidance with VFM-derived pairwise structural biases for coherent spatial propagation, followed by text-conditioned class-wise reasoning. We further introduce Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context. On the global High-Resolution Land Cover (HRLC) benchmark spanning seven datasets across six continents, GeoSeg-OV outperforms the state-of-the-art by +2.5 and +2.7 average mIoU under two training settings. A large-scale zero-shot case study further demonstrates its generalization across geographic domains and category systems without target-domain annotations or retraining.

View source

Similar papers

2026

ConceptSeg: Zero-Shot Geospatial Concept Segmentation for Remote Sensing Images

The existing remote sensing image segmentation methods rely on predefined category labels, which are often insufficient to capture the complex spatial semantics inherent in geospatial concepts, such as flood inundation zones, landslide bodies, and industrial complexes. This letter presents ConceptSeg, a text-guided mul...

Yang Zhao, Ya-Wei Bai, Ming-Ming Jia et al. · 0 citations
Open access 2026

Seeing With Words: Autocaption-Guided Graph Transformer for Remote Sensing Segmentation

Semantic segmentation of remote sensing (RS) imagery is a cornerstone of geospatial analysis, which supports applications, such as land cover mapping, urban planning, and environmental monitoring. Despite significant progress with neural networks, existing approaches remain limited by their reliance on visual features...

Ge Song, Jia-Wei Guo, Yang Zhang et al. · 0 citations
Sep 2026

GeoSeC: Geometric-Semantic Collaborative Learning for Weakly Supervised Remote Sensing Image Segmentation.

Despite the economic advantages of weakly supervised semantic segmentation (WSSS) in remote sensing imagery (RSI), existing VLM- and VFM-based approaches still struggle with domain-specific semantic ambiguities and geometric discontinuities. To address these challenges, we propose GeoSeC, the Geometric-Semantic Collabo...

Xiang-Rong Zhang, Jian-Xun Lai, Guan-Chun Wang et al. · 0 citations
Preprint Sep 2026

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...

Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al. · 0 citations
Jul 2026

StructPointNet: explicit geometric prior extraction and reuse for efficient multimodal remote sensing semantic segmentation

Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, and the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.

Si-Wei Wei, Chao-Jie Wang, Ruo-Xi Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery

Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...

Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.