2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 5638513-5638513· 0 citations· 43 references
Abstract
The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inherent global semantic bias. To address these limitations, we propose the focused adapter (FocA), a plug-and-play parameter-efficient fine-tuning (PEFT) architecture designed to enhance fine-grained perception from frozen VLMs. The FocA features a hybrid structure consisting of two components: an enhancement adapter and an alignment adapter. The enhancement adapter utilizes bottleneck self-attention to capture rich patch-level details often lost during global pooling. The alignment adapter incorporates a cross-modal shared projection subspace to facilitate early feature interaction and implicit alignment. In addition, we develop an explicit shared loss to provide direct semantic supervision, preventing the dilution of critical features during deep propagation. Extensive experiments on the RSICD and RSITMD datasets demonstrate that FocA achieves state-of-the-art (SOTA) performance, attaining a mean recall (mR) of 37.61% on RSICD and 48.95% on RSITMD. Our code is available at https://github.com/WenliangDu/FocA
GLCE (Global-Local Channel Enhancer), a lightweight plug-and-play module with a dual-branch structure, provides a lightweight industrial solution for fine-grained cross-modal retrieval, and offers new insights for channel attention design in Transformer-based vision-language architectures.
Qixuan Pan· International Conference on...· 0 citations
DARAD is proposed, a dual-adapter and ranking-aware distillation framework that preserves the historical cross-modal ranking structure while learning new visual and textual concepts from evolving archives to address the challenge of continual RS-ITR.
Xi Chen, Xu Chen, Xiang-Yang Jia et al.· 0 citations
Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...
Kuang-Rong Hao· International Conference on...· 0 citations
Remote sensing image–text retrieval (RSITR) plays a crucial role in remote sensing interpretation, yet remains challenging under both closed-domain and open-domain scenarios due to semantic noise and domain shifts. To address these challenges within a coherent framework, we propose PriorCLIP, a visual-prior-guided visi...
Jiancheng Pan, Muyuan Ma, Qing Ma et al.· IEEE Transactions on Geoscie...· 12 citations· ⚡1
GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.
Le Yu, Yuan-Wen Wang, Xiao-Tong Qi· IEEE Journal of Selected Top...· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.