Skip to content
Open access

A CNN-transformer multimodal architecture for weakly-supervised audio-visual violence detection

Jul 2026 · Open Access Research Journal of Engineering and Technology · 0 citations

Abstract

Automated detection of violent events in surveillance-scale video is important for timely intervention, but frame-accurate labels are too costly to collect at scale. This has made weakly supervised learning from video-level labels the dominant paradigm. Most existing methods score short video snippets from a single modality and with limited temporal context. We propose a CNN-Transformer multimodal architecture that extracts per-snippet features from frozen, pretrained CNN backbones (I3D for video, VGGish for audio), projects each modality through a lightweight temporal 1D-CNN, and fuses the two modalities across the full temporal extent of a video with a cross-modal Transformer encoder. Snippets are then scored under a multiple-instance-learning (MIL) ranking objective. On XD-Violence, the largest public audio-visual violence detection benchmark, our 4.24M-parameter fusion network reaches 80.14% frame-level average precision (AP) and 93.29% ROC-AUC, outperforming several established baselines while staying much smaller than the CNN backbones it builds on. Ablations show that the audio modality and the Transformer fusion stage each add several points of AP over a visual-only, non-attentive baseline. We also test the model qualitatively on independently sourced, real-world YouTube footage drawn from outside the training distribution. That test exposes a limitation we believe is underappreciated: replacing the benchmark’s undisclosed official visual feature extractor with an independently implemented one collapses the visual predictions, whereas our faithfully reproduced audio pipeline still transfers. Feature-extractor fidelity, and not model architecture alone, is therefore critical for real-world deployment of weakly supervised video anomaly detection systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.