ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering
This work introduces ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs, and develops a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verificatio...