Skip to content
Preprint

TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

TAU-Bench is introduced, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding, and shows that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding.

Abstract

Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. Video anomaly understanding (VAU) seeks to endow models with a similar capability, moving beyond deciding whether a video is anomalous toward explaining how the event develops and why it matters. Although recent vision--language models (VLMs) can generate detailed and plausible anomaly descriptions, their semantic fluency does not ensure that these interpretations remain grounded in the correct anomaly instance over time. Existing benchmarks typically evaluate tracking and semantic understanding through separate protocols, leaving such instance--semantic inconsistency largely unmeasured. We therefore introduce TAU-Bench, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding. TAU-Bench contains 1,118 videos, 1,454 tracks, and 202,438 pixel-level masks spanning 49 event and 45 scene categories, together with track-centric annotations that connect instance-level identification, event-level understanding, and scene-level reasoning. To build TAU-Bench at scale, we developed an automated data engine integrating anomaly suitability filtering, anomaly instance track construction, hierarchical caption annotation, and human quality control. Evaluations across representative VLM families show that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding. These findings therefore highlight instance-grounded evaluation as an important step toward more faithful and reliable VAU systems.

View source

Similar papers

Open access Dec 2024

Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly

Recent progress in video anomaly understanding (VAU) enables significant applications in traffic monitoring and industrial automation, yet existing benchmarks mainly focus on anomaly detection and localization. We push VAU toward practical comprehension by explicitly evaluating whether models can describe what anomaly occurred, explain why it happened, and infer what effect it caused. To this end, we introduce ECVA, a benchmark for Exploring the Causation of Video Anomalies, where each video is paired with detailed human annotations covering (1) anomaly type, temporal boundaries, and event descriptions, (2) natural-language explanations of causes, and (3) free-form descriptions of effects. In addition, ECVA provides annotation-derived importance curves to characterize the relative contribution of key evidence segments, supporting fine-grained evaluation and reliability analysis. Building on ECVA, we propose AnomShield, a video large language model for reasoning-intensive anomaly understanding. AnomShield adopts a Chain-of-Thought reasoning paradigm to explicitly determine and extract anomaly-relevant temporal segments, and subsequently employs spatiotemporal-decoupled positional encoding together with bi-directional state-space modeling to capture fine-grained spatiotemporal dependencies. To evaluate models under ECVA’s setting, we introduce AnomEval, a human-aligned metric for more reliable assessment of video-LLMs. Extensive experiments validate the effectiveness of our benchmark, model, and metric. Code and dataset are available at https://github.com/Dulpy/ECVA.

Hang Du, Guoshun Nan, Jiawen Qian et al. · 9 citations · ⚡3
Preprint Aug 2026

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed, and JeAUG, a metric jointly evaluating semantic interpretability and temporal precision are introduced.

Shibo Gao, Pei-Pei Yang, Xu-Yao Zhang et al. · 0 citations
Preprint Feb 2026

Video-HOCA: A Diagnostic Benchmark for Physical Anomaly Reasoning in Video-LLMs

Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos, is introduced and it is found that recognition outpaces explanation: Task I scores cluster at 75-88, while Task II macro-F1 stays mostly below 50.

Chang Liu, Yunfan Ye, Qingyang Zhou et al. · 0 citations
Preprint Sep 2026

AnomalyCraft-700K: Component-Level Controllable and Verifiable Synthetic Anomalies for Fine-Grained Video Anomaly Understanding

Progress in video anomaly understanding (VAU) has long been limited by inherent deficiencies of real-world anomaly videos, which are hard to collect and offer little control over their content. Synthetic anomaly approaches partially alleviate data scarcity, yet their generation remains largely controlled at the category or prompt level. They also lack component-level verification of video-text consistency and provide insufficient hard normal samples near the normal-anomaly boundary. To address this, we present AnomalyCraft-700K, a component-controllable synthetic anomaly dataset for fine-grained VAU, containing over 40K videos and over 700K task-level textual annotations. From fine-grained semantic components and a progressive three-stage pipeline, we craft anomaly events that are richly detailed, semantically controlled, and temporally structured, and additionally construct per-category hard normal samples to prompt the model to discriminate based on anomaly semantics rather than surface visual cues. Moreover, using the components as verification units, AnomalyCraft-700K further performs component-wise correction of video-text discrepancies introduced during generation, providing reliable annotations with verified cross-modal alignment for six tasks that progress from anomaly detection, through anomaly retrieval and captioning, to fine-grained anomaly reasoning. Evaluations of widely used methods under both traditional and MLLM-based protocols demonstrate that AnomalyCraft-700K serves as an effective source of supervision, from anomaly detection to fine-grained anomaly understanding.

Yu-Zhou Long, Haodong Zhang, Yun-Peng Yang et al. · 0 citations
#computer vision Preprint Aug 2026

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.

Marek Hradil, Danae Sánchez Villegas · 0 citations
Jul 2026

Context-structured Video Anomaly Detection with Large Vision-Language Models

Experiments on UCF-Crime and UBnormal show that CSI-VAD consistently improves over the direct holistic baseline and achieves competitive performance against existing methods, showing the advantage of structured context decomposition for training-free video anomaly detection.

Dongjun Kim, Changjae Oh, Andrea Cavallaro et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.