Skip to content

Focused Adapter: Enhancing Fine-Grained Attention for Remote Sensing Image–Text Retrieval

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5638513-5638513 · 0 citations · 43 references

Abstract

The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inherent global semantic bias. To address these limitations, we propose the focused adapter (FocA), a plug-and-play parameter-efficient fine-tuning (PEFT) architecture designed to enhance fine-grained perception from frozen VLMs. The FocA features a hybrid structure consisting of two components: an enhancement adapter and an alignment adapter. The enhancement adapter utilizes bottleneck self-attention to capture rich patch-level details often lost during global pooling. The alignment adapter incorporates a cross-modal shared projection subspace to facilitate early feature interaction and implicit alignment. In addition, we develop an explicit shared loss to provide direct semantic supervision, preventing the dilution of critical features during deep propagation. Extensive experiments on the RSICD and RSITMD datasets demonstrate that FocA achieves state-of-the-art (SOTA) performance, attaining a mean recall (mR) of 37.61% on RSICD and 48.95% on RSITMD. Our code is available at https://github.com/WenliangDu/FocA

View source

Similar papers

Conference Aug 2026

GLCE: global-local channel enhancer for fine-grained e-commerce image-text retrieval

GLCE (Global-Local Channel Enhancer), a lightweight plug-and-play module with a dual-branch structure, provides a lightweight industrial solution for fine-grained cross-modal retrieval, and offers new insights for channel attention design in Transformer-based vision-language architectures.

Qixuan Pan · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...

Kuang-Rong Hao · 0 citations
May 2024

PriorCLIP: Visual-Prior-Guided Vision–Language Model for Remote Sensing Image–Text Retrieval

Remote sensing image–text retrieval (RSITR) plays a crucial role in remote sensing interpretation, yet remains challenging under both closed-domain and open-domain scenarios due to semantic noise and domain shifts. To address these challenges within a coherent framework, we propose PriorCLIP, a visual-prior-guided visi...

Jiancheng Pan, Muyuan Ma, Qing Ma et al. · 12 citations · ⚡1
Open access 2026

GeoVP: A Unified Visual Prompting Framework for Multisource Remote-Sensing Image Understanding

GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding, enabling unified image-level and region-level understanding under different prompt granularities and demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.

Le Yu, Yuan-Wen Wang, Xiao-Tong Qi · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.