Skip to content

ConceptSeg: Zero-Shot Geospatial Concept Segmentation for Remote Sensing Images

2026 · IEEE Geoscience and Remote Sensing Letters · Vol 23, pp. 6019805-6019805 · 0 citations · 19 references

Abstract

The existing remote sensing image segmentation methods rely on predefined category labels, which are often insufficient to capture the complex spatial semantics inherent in geospatial concepts, such as flood inundation zones, landslide bodies, and industrial complexes. This letter presents ConceptSeg, a text-guided multimodal segmentation framework that enables zero-shot segmentation of linguistically described geospatial objects in remote sensing imagery. ConceptSeg employs a decoupled architecture comprising a multiscale image encoder with 2-D rotary position embedding (2D-RoPE), a unified multimodal prompt encoder with adaptive gated fusion, and a compact two-way cross-attention decoder, bridging natural-language concept descriptions with pixel-level predictions through region of interest (ROI)-guided feature extraction and contrastive alignment. With only $\approx $ 35 M trainable parameters atop frozen encoders, it stays markedly lighter than recent multimodal large-language-model baselines. Experiments on multiple benchmarks show up to +0.09 mIoU improvement over prior cascaded reasoning baselines out-of-domain, a smaller margin over concurrent open-vocabulary methods that is significant only with an external box prior, and consistent gIoU/cIoU gains in-domain, supporting concept-driven segmentation for complex geospatial reasoning.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.