Skip to content
Open access

Adapting Pure-Text Large Language Models for Remote Sensing Classification via Lightweight Visual Adapter

Jul 2026 · Applied and Computational Engineering · Vol 257, pp. 14-20 · 0 citations

TL;DR

External trainable mapping subnetworks, as the test records suggest, unlock visual discrimination capacity for unmodified text-only large language models, supplying a low-hardware threshold tuning route for earth observation research groups constrained by computing resources.

Abstract

Stable cross-domain feature alignment is indispensable for earth observation classification, which is fundamentally hampered by radiometric gaps between generic pre-training images and aerial remote sensing data. Vision-language pre-trained models exhibit strong zero-shot capability on ordinary photos yet suffer severe accuracy loss on satellite and aerial imagery. While LoRA tuning cuts partial training costs, backbone parameter fine-tuning still brings considerable GPU memory overhead in training. Relying on frozen DeepSeek V4 MoE text LLM and static SigLIP vision encoder, this study designs a slim cross-modal projection subnet to eliminate feature distribution gaps between modalities. Stacked residual MLPs constitute the sole learnable part, containing roughly 20M parameters for visual-text latent space matching. The model is evaluated collectively on EuroSAT, PatternNet and RSSCN7, covering nearly 60,000 aerial images with 55 separate scene classes. Recorded aggregate classification precision reached 99.80% across the unified multi-source testing pool. Compared with LoRA-dependent VL-ZSDA-RS benchmark schemes, the adjustable parameter scale shrinks by over half, alongside a 46% cut in peak GPU memory usage. Layer-wise ablation trials reflect unstable matching performance under shallow projection layouts; five stacked transformation layers deliver the most balanced tradeoff between computation overhead and inter-modal alignment quality. External trainable mapping subnetworks, as the test records suggest, unlock visual discrimination capacity for unmodified text-only large language models, supplying a low-hardware threshold tuning route for earth observation research groups constrained by computing resources.

Read PDF

Similar papers

2026

Toward Zero-Forgetting: A Training-Free Multimodal Framework for Remote Sensing Class-Incremental Learning

Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from m...

Wen-Liang Du, Ji-Cun He, Jia-Qi Zhao et al. · 0 citations
Preprint Sep 2026

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

General-purpose vision-language models (VLMs) now support strong visual recognition, instruction following, and generation. However, most pretrained visual encoders are built around three-channel natural images and do not directly accommodate observations such as native multispectral measurements or synthetic aperture...

Shan-Ji Liu, Ke-Lu Yao, Jun-Xiao Xue et al. · 0 citations
2026

Focused Adapter: Enhancing Fine-Grained Attention for Remote Sensing Image–Text Retrieval

The emergence of large-scale vision–language models (VLMs) has significantly advanced remote sensing image–text retrieval (RSITR) by providing powerful cross-modal semantic priors. However, when adapted to the remote sensing (RS) domain, these models struggle to capture fine-grained representations due to their inheren...

Wen-Liang Du, Xiao-Yu Xu, Jia-Qi Zhao et al. · 0 citations
Open access Jul 2026

Meta-Prompting with Open-Source Language Models for Zero-Shot Scene Classification in Remote Sensing

This paper investigates whether meta-prompting with large language models (LLMs) can improve zero-shot scene classification in RS by automatically generating semantically rich class descriptions and highlights the potential of open-source LLMs as scalable prompt generators for zero-shot remote-sensing recognition.

Antonis Promponas, Eirini Baltzi, Valsamis Ntouskos et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data

Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imagi...

Xian-Yang Miao, Ke-Lu Yao, Ye-Hua Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.