Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Amateur team sport is a useful, largely untested place to check that assumption: over eight million students played a school sport in the United States in 2024-25 alone, almost none of it filmed by more than a single fixed camera, with several candidate actors crowded into frame and no operator or second angle to fall back on. Using volleyball as a test case, we ask whether strong performance on general video and world-model benchmarks translates into reliable, per-player attribution once footage is this chaotic, turning footage into statistics through a chain of tasks from finding play boundaries to naming who did what. We evaluate four approaches (prompting and agentic reasoning over frontier vision-language models, classical computer vision with small trained specialists, self-supervised video world models, and manual annotation) at every stage, on 66 amateur matches with 46,648 human-labelled contacts, filmed under conditions no published benchmark uses. No single paradigm wins every stage, and static, single-frame computer vision is not competitive at any stage involving motion or identity. A prompted model segments matches well, yet a far smaller trained model beats it at spotting contacts for a fraction of the cost, and the sport's own rules recover rally outcomes the pixels cannot. Identity is where every automated approach struggles: a jersey number is a static fact temporal reasoning cannot recover if never visible, unlike sporting action, a repeated motor pattern a temporal model can exploit, which is why holistic reasoning improves event detection while identity stays unchanged. We close with where each approach earns its cost, and what transfers beyond volleyball to amateur sport.
A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers...
En-Xin Song, Yi-Nuo Xu, Shu-Sheng Yang et al.· 0 citations
MultiVENT-Raw is released, a multilingual collection of nearly 120,000 primarily raw videos paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos.
Reno Kriz, David Etter, Alexander Martin et al.· 0 citations
VideoVIBE is introduced, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks and V2Lens is proposed, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-...
Jia-Jun Xu, Yang-Hao Zhou, Jing Liao et al.· 1 citation
Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal continuity of game events nor support complex tasks such as action anticipation. To address this issue,...
Yunhao Zhao, Haoying Sun, Jiarui Li et al.· 0 citations
MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion, the variables a world model must predict, is introduced, a pair of near-identical clips that differ only in motion.
Dhairya Bhatia, Bishoy M. Galoaa, Oliver Fritsche et al.· 0 citations
CinematicVQA is introduced, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing the introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrati...
Shuo Xing, Pooja Verlani, Balu Adsumilli et al.· 0 citations
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.