Skip to content

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Dazhao Du Jian Liu Jialong Qin Tao Han Bohai Gu Fangqi Zhu Yujia Zhang Eric Liu Xi Chen Song Guo
Sep 2026
Artificial Intelligence Computer Vision

Abstract

Video large language models (Video LLMs) can achieve strong video-QA accuracy without reliably tracking spatiotemporal dynamics. A model may answer a motion question from static cues, for example, and give the same prediction even after the underlying motion is reversed. Correctness-based reinforcement learning does not directly address this problem because it rewards the final answer without requiring the policy to respond to task-relevant video dynamics. We propose Counterfactual Relational Policy Optimization (CRPO), which explicitly trains Video LLMs to respond to controlled changes in the visual input. For each training example, CRPO constructs a counterfactual video, such as a horizontally flipped or temporally reversed version of the original, and jointly optimizes rollouts from both videos under a shared policy. Factual supervision anchors what the model should answer, while counterfactual supervision constrains when that answer should change: predictions should change when an intervention alters task-relevant dynamics and remain stable when the queried property is preserved. This coupling provides a direct behavioral learning signal without requiring ground-truth labels for transformed videos or annotated spatiotemporal reasoning traces, while discouraging indiscriminate answer changes. To evaluate this property, we introduce DyBench, a paired counterfactual benchmark with 3{,}014 videos and a strict pair-accuracy metric. Across paired spatiotemporal evaluations and standard video benchmarks, CRPO improves sensitivity to motion and temporal changes while improving performance on general video understanding. The gains also extend to segment reordering, a transformation never used during training, suggesting that CRPO learns sensitivity to video dynamics beyond the training interventions. The project website can be found at https://ddz16.github.io/crpo.github.io/ .

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Diffusion models as plug-and-play priors

The possibility of inferring high-dimensional data inference in a model that consists of a prior and an auxiliary differentiable constraint given some additional information is considered, thereby allowing a range of potential applications in adapting models to new domains and tasks.

Alexandros Graikos, Esmeralda S. Whitammer, N. Jojic et al. · 316 citations · ⚡15

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.