PriorCLIP: Visual-Prior-Guided Vision–Language Model for Remote Sensing Image–Text Retrieval
Abstract
Remote sensing image–text retrieval (RSITR) plays a crucial role in remote sensing interpretation, yet remains challenging under both closed-domain and open-domain scenarios due to semantic noise and domain shifts. To address these challenges within a coherent framework, we propose PriorCLIP, a visual-prior-guided vision–language model that organizes closed-domain and open-domain retrieval as two complementary realizations of the same learning principle. PriorCLIP first introduces remote sensing scene knowledge as a visual prior, then uses this prior to adapt image and text representations to the available training regime, and finally regularizes the representations in a discriminative common space. In closed-domain retrieval, this principle is instantiated at the sample level: the spatial progressive attention encoder (S-PAE) constructs an instruction-conditioned belief matrix to reweight visual tokens and suppress noisy regions, while the temporal progressive attention encoder (T-PAE) progressively propagates contextual information to strengthen text representations. In open-domain retrieval, the same principle is extended to representation transfer through a two-stage strategy that first pretrains image and text encoders on large-scale coarse-grained pairs and then performs vision-instruction fine-tuning on fine-grained pairs. A cluster-based symmetric contrastive affiliation loss further exploits scene-category structure to improve intraclass compactness and interclass separation, thereby reducing semantic confusion in the common space. Extensive experiments on RSICD and RSITMD benchmarks demonstrate that PriorCLIP achieves substantial improvements, outperforming existing methods by 4.9% and 4.0% in closed-domain retrieval, and by 7.3% and 9.4% in open-domain retrieval, respectively.