Jul 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 237-247· 0 citations
TL;DR
Results demonstrate that combining anomaly-aware attention with keyframe extraction in a single trainable pipeline yields both higher detection fidelity and immediately actionable, human-reviewable video summaries, making the framework well suited for real-time deployment in smart-city and public-safety surveillance infrastructures.
Abstract
The exponential growth of closed-circuit television (CCTV) infrastructure has far outpaced the capacity of human operators to monitor footage in real time, leaving a critical gap between data acquisition and actionable situational awareness. This paper proposes a unified deep neural network framework that jointly performs automated crime (anomalous-event) detection and attention-guided video summarization from untrimmed surveillance streams. The framework couples a 3D convolutional feature extractor with a bidirectional long short-term memory (BiLSTM) temporal encoder and a self-attention module to model both short-range spatial cues and long-range temporal dependencies characteristic of criminal activities such as assault, robbery, arson, and shooting. The learned frame-level attention scores are re-used, without additional supervision, to drive a diversity-aware keyframe selection module that automatically compresses hours of footage into a compact, timestamped evidentiary summary whenever an anomalous segment is detected. The proposed system was evaluated on the UCF-Crime and DCSASS benchmark datasets for detection and on TVSum/SumMe-style protocols for summarization quality. Experimental results indicate that the framework attains a detection accuracy of 98.4%, an F1-score of 0.978, and an AUC of 0.991, outperforming C3D, two-stream CNN, and CNN-LSTM baselines by 3.4-16.0 percentage points, while the attached summarization module achieves an F-score of 0.71 with a video compression ratio exceeding 92% and end-to-end inference throughput of 41 frames per second on a single consumer-grade GPU. These results demonstrate that combining anomaly-aware attention with keyframe extraction in a single trainable pipeline yields both higher detection fidelity and immediately actionable, human-reviewable video summaries, making the framework well suited for real-time deployment in smart-city and public-safety surveillance infrastructures.
A deep learning-based forensic framework for real-time detection of suspicious human activity in CCTV videos, trained without relying on any external sensors is proposed, and incorporates anonymization of personal identities and local edge-based processing to prevent raw data exposure.
Qazi Mazhar Ul Haq, Muhammad Imran, M. Waqas et al.· Arab Journal of Forensic Sci...· 0 citations
Automated detection of violent events in surveillance-scale video is important for timely intervention, but frame-accurate labels are too costly to collect at scale. This has made weakly supervised learning from video-level labels the dominant paradigm. Most existing methods score short video snippets from a single mod...
Rahul Gyawali, Rabin Dhakal, Deepak Dulal et al.· Open Access Research Journal...· 0 citations
A hybrid deep learning (DL) model that integrates Coordinate Attention (CA), Convolutional Neural Networks (CNN), and Transformer encoders for better HAR achieves superior performance and proved its effectiveness for workplace safety monitoring applications.
E. Jyotsna, T. Jarin· Neural computing & applicati...· 0 citations
This work proposes a novel taxonomy categorizing 34 existing language-driven VAD methods based on learning paradigm, core architecture, adaptation strategy, and functional output, and serves as a foundational resource for the VAD community, advancing the understanding and application of language-driven models in addres...
Olfa Saket, A. Ben Aicha, H. Fathallah· International Journal of Mul...· 0 citations
Violence recognition is an urgent need for real-time human activity identification in surveillance video streams, which is becoming more important in areas including public safety, law enforcement, and security monitoring. Even while visual understanding has come a long way, finding a balance between accuracy and spe...