Skip to content
Preprint

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

Sep 2026 · 0 citations · 24 references
Computer Science

TL;DR

CinematicVQA is introduced, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing the introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions.

Abstract

Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.

View source

Similar papers

#natural language process... Preprint Aug 2026

DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.

Bomiao Wang, Zekai Shao, Jiexiang Lan et al. · 0 citations
Preprint Aug 2026

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

This work introduces TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent, and establishes directorial intent as a previously overlooked dimension of multimodal understanding.

Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant et al. · 0 citations
Preprint Sep 2026

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time, leading all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

Yu-Bo Zhu, Ya-Wen Shao, Zi-Yun Dai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles

Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce Ci...

Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Aug 2026

A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding

A3Bench is introduced, an audience-aligned benchmark for evaluating video audience insights with large-scale videos and high-quality multilingual comments and Cognition Interaction of Thought (CIoT), a structured reasoning framework that emulates key aspects of cognitive processes is proposed.

Yi-Ming Lei, Guozhen Peng, Ze-Ming Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.