Jul 2026· The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences· 1 citation· 9 references
TL;DR
This paper investigates whether meta-prompting with large language models (LLMs) can improve zero-shot scene classification in RS by automatically generating semantically rich class descriptions and highlights the potential of open-source LLMs as scalable prompt generators for zero-shot remote-sensing recognition.
Abstract
Abstract. Zero-shot visual recognition with vision-language models (VLMs) has shown strong generalization to unseen categories in natural-image benchmarks, yet its effectiveness in remote-sensing (RS) imagery remains less explored. In this paper, we investigate whether meta-prompting with large language models (LLMs) can improve zero-shot scene classification in RS by automatically generating semantically rich class descriptions. Building on the Meta-Prompting for Visual Recognition (MPVR) framework, we evaluate three open-source LLMs, Mixtral-8×7B, Qwen 2.5 7B, and LLaMA 3.1 8B, as prompt generators across five RS benchmark datasets. The resulting descriptions are encoded with several VLMs, including CLIP, MetaCLIP, RemoteCLIP, and CLIP-LAION-RS, and compared against generic single-template and handcrafted domain-specific prompting baselines. Our results show that LLM-generated prompts are competitive with, and in several cases improve upon, manually designed templates, while revealing that the gains depend on both the dataset and the visual backbone. Overall, the study highlights the potential of open-source LLMs as scalable prompt generators for zero-shot remote-sensing recognition and provides insight into the transferability of meta-prompting beyond natural-image domains.
Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from m...
Wen-Liang Du, Ji-Cun He, Jia-Qi Zhao et al.· IEEE Transactions on Geoscie...· 0 citations
Cross-scene remote sensing (RS) image classification is crucial for urban analysis and environmental monitoring, but often suffers from distribution shifts caused by different imaging conditions. While unsupervised domain adaptation (UDA) can alleviate cross-scene discrepancies, its closed-set assumption is violated wh...
Xin Zhao, Yuanyuan Ye, Jun Lin et al.· IEEE Transactions on Geoscie...· 0 citations
A systematic survey and diagnostic evaluation of MLLMs for RSISU and demonstrates the strong transferability of general-purpose CV-MLLMs and shows that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks.
The AP-LCA approach introduces a novel local cross-modal alignment strategy that utilizes image cropping and similarity-based semantic contribution assessment to precisely map fine-grained descriptions to relevant local image regions.
Si-Ying Wu, Song Wu· International Conference on...· 0 citations
UC-VLM is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted pro...
Lei Tan, Shuwei Li, Mohan S. Kankanhalli et al.· 0 citations
Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly complex designs such as heavy cross-modal fusion, staged training, and iterative annotation pipelines. We revisit whether such complexity is nec...
Xiao-Yan Wei, Zhi-Min Yao, Rui-Lin Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.