Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos, is introduced and it is found that recognition outpaces explanation: Task I scores cluster at 75-88, while Task II macro-F1 stays mostly below 50.
Abstract
We introduce Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos. Video-HOCA uses an Ontological-Causal taxonomy to distinguish violations of an entity's own properties or capabilities from violations of physical relations among entities and the environment. It contains more than 1,400 generated and real-world videos and 3,470 question-answer pairs, with human verification of labels and reference answers. The benchmark evaluates four levels of reasoning: plausibility checking, anomaly attribution, fine-grained recognition, and open-ended physical explanation. Across 20 Instruct-mode Video-LLMs, we find that recognition outpaces explanation: Task I scores cluster at 75-88, while Task II macro-F1 stays mostly below 50. We also find that the Ontological-Causal gap depends on the task and model configuration, and that Thinking-mode gains are not explained by sampling or output budget alone. Annotation agreement, Task-IV human-judge and judge-judge checks, alternative metrics, and temporal/decoding controls validate the evaluation pipeline and bound the claims supported by the benchmark.
TAU-Bench is introduced, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding, and shows that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding.
Kepeng Yang, Dong-Xuan Liu, Rongxin Gao et al.· 0 citations
Recent progress in video anomaly understanding (VAU) enables significant applications in traffic monitoring and industrial automation, yet existing benchmarks mainly focus on anomaly detection and localization. We push VAU toward practical comprehension by explicitly evaluating whether models can describe what anomaly occurred, explain why it happened, and infer what effect it caused. To this end, we introduce ECVA, a benchmark for Exploring the Causation of Video Anomalies, where each video is paired with detailed human annotations covering (1) anomaly type, temporal boundaries, and event descriptions, (2) natural-language explanations of causes, and (3) free-form descriptions of effects. In addition, ECVA provides annotation-derived importance curves to characterize the relative contribution of key evidence segments, supporting fine-grained evaluation and reliability analysis. Building on ECVA, we propose AnomShield, a video large language model for reasoning-intensive anomaly understanding. AnomShield adopts a Chain-of-Thought reasoning paradigm to explicitly determine and extract anomaly-relevant temporal segments, and subsequently employs spatiotemporal-decoupled positional encoding together with bi-directional state-space modeling to capture fine-grained spatiotemporal dependencies. To evaluate models under ECVA’s setting, we introduce AnomEval, a human-aligned metric for more reliable assessment of video-LLMs. Extensive experiments validate the effectiveness of our benchmark, model, and metric. Code and dataset are available at https://github.com/Dulpy/ECVA.
Hang Du, Guoshun Nan, Jiawen Qian et al.· International Journal of Com...· 9 citations· ⚡3
Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed, and JeAUG, a metric jointly evaluating semantic interpretability and temporal precision are introduced.
Shibo Gao, Pei-Pei Yang, Xu-Yao Zhang et al.· 0 citations
CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios, and all models degrade sharply on accident type and responsibility reasoning.
Sparsh Garg, Yi-Wen Chen, Vijay Kumar et al.· 0 citations
Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often conflating visual quality with cognitive correctness. We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks. To enable precise diagnosis, we design a three-level VLM-as-Judge protocol that independently assesses video-level fluency, task-level rule adherence, and sample-level goal realization. Evaluations of leading models reveal a striking gap: while models achieve strong rendering scores, they consistently fail on logic-heavy and rule-constrained tasks. To address this, we propose Vid-PRE (Video Prompt Reasoner and Enhancer), a model-agnostic prompt rewriter that offloads the cognitive burden of reasoning to a dedicated VLM. Trained via reinforcement learning with purely text-based rewards, Vid-PRE produces concise, constraint-aware prompts without the instability of video-level reward signals. Experiments show that Vid-PRE yields substantial reasoning improvements across multiple generators without architectural modifications. Together, VWG-Bench and Vid-PRE offer a rigorous diagnostic lens and a scalable path toward true think-with-video capabilities. All data and code are publicly available at https://huggingface.co/datasets/KlingTeam/VWG-Bench.
Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al.· 0 citations
We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection. TAR contains 44,040 chain-of-thought training annotations across 10 tasks for 3,670 CCTV videos ($\sim$26 hours) from eight public datasets. Its evaluation component, TAR-Bench, contains 960 human-curated test annotations for 80 held-out clips trimmed from 17 public YouTube videos. TAR's training annotations are produced with MAVEN, which consolidates multi-scale video evidence into structured event descriptions before generating question-answer pairs and reasoning traces. On TAR-Bench, eleven vision-language models reveal that strong question-answering accuracy does not reliably predict temporal or scene reasoning ability. Multi-task fine-tuning on TAR yields consistent gains, with the full 10-task model improving aggregate score by 21.4 points over its zero-shot baseline. TAR and TAR-Bench provide the official training and in-domain evaluation data for AI City Challenge 2026 Track 3. The dataset is available at https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning
Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.