This work proposes TAU-Agent, an agentic retrieval-augmented framework for traffic anomaly understanding that orchestrates two visual perception tools to retrieve and select query-relevant evidence, including captions, temporal intervals, and object trajectories.
Abstract
Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agentic retrieval-augmented framework for traffic anomaly understanding. Given a task query, a central retrieval agent orchestrates two visual perception tools, namely a Video Captioning Tool and an Open-Vocabulary Tracking Tool, to retrieve and select query-relevant evidence, including captions, temporal intervals, and object trajectories. The selected evidence, together with sampled video frames and the input query, is provided to a supervised fine-tuned vision-language model for final reasoning and answer generation. We evaluate TAU-Agent on both the in-domain and the out-of-domain benchmarks from the AI City Challenge 2026. TAU-Agent achieves scores of 0.6779 on Track 3, 0.3998 on Track 7, and 67.9275 on Track 8, ranking second, twelfth, and fifth, respectively. Code is available at: https://github.com/siri-rouser/TAU-Agent.
UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.
Peng Li, Qianqian Xu, Shilong Bao et al.· 1 citation
TAR and TAR-Bench serve as the official training and in-domain evaluation resources for AI City Challenge 2026 Track 3 and highlight persistent limitations in temporal precision and causal attribution.
Han Zhang, Yi-Lin Zhao, Zaid Pervaiz Bhat et al.· 1 citation
This survey reviews video understanding agents: systems that use video as the primary information source and solve understanding tasks through adaptive state construction and action selection, and identifies open directions toward agentic-native temporal modeling and video-native agents.
Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed, and JeAUG, a metric jointly evaluating semantic interpretability and temporal precision are introduced.
Shibo Gao, Pei-Pei Yang, Xu-Yao Zhang et al.· 0 citations
ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals, is introduced andlations show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents provide the strongest trajectory-relative su...
Hao-Dong Chen, Shuai Wang, Yu Yin et al.· 0 citations
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but of...
Mohd Ubaid Wani, Sara Atito, Josef Kittler et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.