CinematicVQA is introduced, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing the introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions.
Abstract
Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.
DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.
Bomiao Wang, Zekai Shao, Jiexiang Lan et al.· 0 citations
This work introduces TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent, and establishes directorial intent as a previously overlooked dimension of multimodal understanding.
Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant et al.· 0 citations
WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time, leading all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.
Yu-Bo Zhu, Ya-Wen Shao, Zi-Yun Dai et al.· 0 citations
Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce Ci...
Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei· 0 citations
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...
A3Bench is introduced, an audience-aligned benchmark for evaluating video audience insights with large-scale videos and high-quality multilingual comments and Cognition Interaction of Thought (CIoT), a structured reasoning framework that emulates key aspects of cognitive processes is proposed.
Yi-Ming Lei, Guozhen Peng, Ze-Ming Liu et al.· Frontiers of Computer Scienc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.