Skip to content

What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation

Jul 2026 · arXiv.org · Vol abs/2607.12304 · 0 citations · 20 references
Computer Science

TL;DR

A label-free screen for the channel question, the reversal-drop: the accuracy lost when the visual sequence is reversed while RoPE remains forward, which can be applied to compatible temporal benchmarks without new annotations.

Abstract

A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1. The task question: is the question even temporal, does it need several frames and their order? and 2. The channel question, when it does, does the model recover the order from the pixels, or read it off the positional encoding (RoPE)? Most of a temporal score answers neither, a single frame and answer priors often carry it. The field's validity checks, frame-shuffle sensitivity and the accuracy gained from the full video, speak only to the task question. We contribute a label-free screen for the channel question, the reversal-drop: the accuracy lost when the visual sequence is reversed while RoPE remains forward. It can be applied to compatible temporal benchmarks without new annotations. Paired reverse labels, or tasks whose labels transform deterministically under reversal, distinguish models that follow reversed content from those merely disrupted by the conflict. Molmo2 answers the forward event reading order off positions, while Qwen3-VL answers the reversed event it actually sees, reading visual order (comparatively). We call them position-dominant and visual-sequence-dominant. The split holds across two benchmarks and several temporal tasks at two scales, and activation patching shows it is a real internal property, not an artifact of the conflict. The distinction matters, the two channels fail on opposite inputs so two models with similar score are not interchangable, i.e. an aggregate score does not reflect potential failure modes.

View source

Similar papers

Open access 2026

Speaking ITM’s Language: Query Reformulation and Temporal-MMR for Long-Video QA

The resulting training-free pipeline, RECAST, consistently outperforms recent frame-selection baselines without modifying the ITM encoder, and preserves the per-frame matching cost of a standard single-query baseline.

S. Han, Thang Vu, Junyeong Kim · 0 citations
#computer vision Preprint Aug 2026

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.

Marek Hradil, Danae Sánchez Villegas · 0 citations
Preprint Sep 2026

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion, the variables a world model must predict, is introduced, a pair of near-identical clips that differ only in motion.

Dhairya Bhatia, Bishoy M. Galoaa, Oliver Fritsche et al. · 0 citations
Preprint Aug 2026

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

VES-Bench, a 600-question benchmark of Temporal Ordering and Event Counting items over 348 public long videos, is introduced and TRACE, a training-free agent that grounds answers in raw visual clips, builds an evidence bundle round by round, and stops only when the answer stabilises as the bundle grows and a final pass...

Peng Liu, Jun-Bo Niu, Xiaoyan Hu et al. · 2 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.