GeoSeg-OV is proposed, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding, and introduces Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context.
Abstract
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization. Recent attempts have begun to incorporate auxiliary vision foundation models (VFMs), typically coupling their features with text embeddings as additional matching evidence. However, this strategy may introduce inconsistent matching signals while leaving the structure-sensitive representations of VFMs insufficiently exploited. We therefore propose GeoSeg-OV, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding. GeoSeg-OV constructs an orientation-robust cost volume from multi-rotation CLIP features, while a frozen VFM extracts multi-scale structure-sensitive features in parallel. We propose Structure-Guided Aggregation (SGA), which integrates cost tokens and CLIP semantic guidance with VFM-derived pairwise structural biases for coherent spatial propagation, followed by text-conditioned class-wise reasoning. We further introduce Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context. On the global High-Resolution Land Cover (HRLC) benchmark spanning seven datasets across six continents, GeoSeg-OV outperforms the state-of-the-art by +2.5 and +2.7 average mIoU under two training settings. A large-scale zero-shot case study further demonstrates its generalization across geographic domains and category systems without target-domain annotations or retraining.
The existing remote sensing image segmentation methods rely on predefined category labels, which are often insufficient to capture the complex spatial semantics inherent in geospatial concepts, such as flood inundation zones, landslide bodies, and industrial complexes. This letter presents ConceptSeg, a text-guided mul...
Yang Zhao, Ya-Wei Bai, Ming-Ming Jia et al.· IEEE Geoscience and Remote S...· 0 citations
Semantic segmentation of remote sensing (RS) imagery is a cornerstone of geospatial analysis, which supports applications, such as land cover mapping, urban planning, and environmental monitoring. Despite significant progress with neural networks, existing approaches remain limited by their reliance on visual features...
Ge Song, Jia-Wei Guo, Yang Zhang et al.· IEEE Journal of Selected Top...· 0 citations
Despite the economic advantages of weakly supervised semantic segmentation (WSSS) in remote sensing imagery (RSI), existing VLM- and VFM-based approaches still struggle with domain-specific semantic ambiguities and geometric discontinuities. To address these challenges, we propose GeoSeC, the Geometric-Semantic Collabo...
Xiang-Rong Zhang, Jian-Xun Lai, Guan-Chun Wang et al.· IEEE Transactions on Image P...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, and the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.
Si-Wei Wei, Chao-Jie Wang, Ruo-Xi Wang et al.· Journal of Applied Remote Se...· 0 citations
Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...
Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.