Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns
This paper investigates LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, and proposes VAD-DPO, a direct preference optimization method supervised with counter-example pairs that enhances both anomaly detection and reasoning performance, particularly in scene-dependent scenarios.