Skip to content
Preprint

ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

Sep 2026 · 2 citations · 56 references
Computer Science

TL;DR

This work introduces ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs, and develops a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA

View source

Similar papers

Preprint Sep 2026

Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing

A selective tool use framework in which a single VLM either answers directly or invokes a deterministic change analysis tool to obtain question specific evidence to demonstrate the benefit of question-specific semantic evidence for Change VQA, while highlighting the influence of semantic prediction quality on the resul...

Y. Bazi, M. M. Al Rahhal, M. Mekhtiche et al. · 0 citations
Preprint Oct 2026

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons...

Leon D. Mayer, Lucas Luttner, P. Godau et al. · 0 citations
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
2026

CRISP: Cross-Modal Residual Guidance and Spatial Realignment for Remote Sensing Visual Question Answering

Remote sensing visual question answering (RSVQA) remains challenging because the evidence relevant to the question in remote sensing (RS) imagery is typically sparse, spatially dispersed, and highly variable in scale and layout. This structural heterogeneity poses a major challenge to existing transfer strategies, whic...

Chang Xu, Zhong-Le Ren, Biao Hou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.