Skip to content
Preprint

SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

It is shown that higher resolution is not universally beneficial and, once sufficient granularity is reached, token quality matters more than token quantity, and a plug-and-play framework that unifies task-adaptive token granularity allocation with holistic geospatial token importance modulation is presented.

Abstract

RS-LVLMs have advanced multimodal understanding of Earth observation imagery, yet their performance is fundamentally constrained by high-resolution processing, as visual token counts grow quadratically with linear input resolution while important visual evidence is inherently sparse and increasingly diluted across the expanded sequence. Existing token pruning methods largely rely on scale-agnostic resolution policies and isolated importance cues, limiting task-aligned granularity adaptation and holistic evidence preservation. To address this, we present Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning (SA-GEM), a plug-and-play framework that unifies task-adaptive token granularity allocation with holistic geospatial token importance modulation. Specifically, a lightweight router selects the resolution based on query-dependent token granularity, while a token importance modulator jointly models task relevance, spatial structure, and local redundancy to preserve holistic geospatial evidence. We show that higher resolution is not universally beneficial and, once sufficient granularity is reached, token quality matters more than token quantity. Experiments across various benchmarks demonstrate that SA-GEM achieves consistent gains in both accuracy and efficiency over existing pruning methods. On XLRS-Bench, it surpasses GeoLLaVA-8K by 2.3% in accuracy with a 2.4 times total inference speedup.

View source

Similar papers

Open access Aug 2026

GeoGATE: Geo-Sensor-Guided Adaptive Token and Evidence Reasoning for High-Resolution Remote Sensing Image Understanding

GeoGATE is introduced, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning that associates adaptive slicing most strongly with localization, retrieval with language and...

Jing-Nan Zhang, Feng-Jun Zhang · 0 citations
Preprint Aug 2026

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

GeoSeg-OV is proposed, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding, and introduces Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context.

Rui-Zhong Liu, Ting-Zhang Luo, Zai-Yan Zhang et al. · 0 citations
Open access Aug 2026

TARA: Task-Adaptive Rank Allocation for Efficient Large Language Model Fine-Tuning in Geo-Information Text Classification

Geo-information texts, including geospatial data-use regulations and Earth observation metadata, are central to data governance and compliance auditing in remote sensing ecosystems. Full fine-tuning of large pre-trained language models is often computationally impractical, while standard LoRA reduces cost but assigns a...

Can-Hui Wang, Juntao Shen, Yi-Cong Feng et al. · 0 citations
Preprint Aug 2026

AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language by using a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network.

Md Aminur Hossain, Omkumar Vaghasiya, R. Dwivedi et al. · 0 citations
Preprint Aug 2026

CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the...

Yutian Jiang, Jiabo Liu, Xixuan Hao et al. · 1 citation
Open access Dec 2025

RAH-VLA: Resolution-Adaptive Hierarchical Vision–Language Alignment for Multimodal Remote Sensing Understanding

Multimodal vision–language modeling has emerged as a promising paradigm for remote sensing (RS) image understanding. However, existing methods are limited by fixed-resolution visual processing and single-scale vision–language alignment, making it difficult to simultaneously preserve fine-grained details and maintain se...

Siyu Zhang, Lianlei Shan, Runhe Qiu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.