Skip to content

PriorCLIP: Visual-Prior-Guided Vision–Language Model for Remote Sensing Image–Text Retrieval

May 2024 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5639516-5639516 · 12 citations · ⚡ 1 influential · 94 references
Computer Science

Abstract

Remote sensing image–text retrieval (RSITR) plays a crucial role in remote sensing interpretation, yet remains challenging under both closed-domain and open-domain scenarios due to semantic noise and domain shifts. To address these challenges within a coherent framework, we propose PriorCLIP, a visual-prior-guided vision–language model that organizes closed-domain and open-domain retrieval as two complementary realizations of the same learning principle. PriorCLIP first introduces remote sensing scene knowledge as a visual prior, then uses this prior to adapt image and text representations to the available training regime, and finally regularizes the representations in a discriminative common space. In closed-domain retrieval, this principle is instantiated at the sample level: the spatial progressive attention encoder (S-PAE) constructs an instruction-conditioned belief matrix to reweight visual tokens and suppress noisy regions, while the temporal progressive attention encoder (T-PAE) progressively propagates contextual information to strengthen text representations. In open-domain retrieval, the same principle is extended to representation transfer through a two-stage strategy that first pretrains image and text encoders on large-scale coarse-grained pairs and then performs vision-instruction fine-tuning on fine-grained pairs. A cluster-based symmetric contrastive affiliation loss further exploits scene-category structure to improve intraclass compactness and interclass separation, thereby reducing semantic confusion in the common space. Extensive experiments on RSICD and RSITMD benchmarks demonstrate that PriorCLIP achieves substantial improvements, outperforming existing methods by 4.9% and 4.0% in closed-domain retrieval, and by 7.3% and 9.4% in open-domain retrieval, respectively.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.