Skip to content
Review Open access

A Deep Neural Network Framework for Automated Crime Detection and Video Summarization

Jul 2026 · International journal of computer information systems and industrial management applications · Vol 18, pp. 237-247 · 0 citations

TL;DR

Results demonstrate that combining anomaly-aware attention with keyframe extraction in a single trainable pipeline yields both higher detection fidelity and immediately actionable, human-reviewable video summaries, making the framework well suited for real-time deployment in smart-city and public-safety surveillance infrastructures.

Abstract

The exponential growth of closed-circuit television (CCTV) infrastructure has far outpaced the capacity of human operators to monitor footage in real time, leaving a critical gap between data acquisition and actionable situational awareness. This paper proposes a unified deep neural network framework that jointly performs automated crime (anomalous-event) detection and attention-guided video summarization from untrimmed surveillance streams. The framework couples a 3D convolutional feature extractor with a bidirectional long short-term memory (BiLSTM) temporal encoder and a self-attention module to model both short-range spatial cues and long-range temporal dependencies characteristic of criminal activities such as assault, robbery, arson, and shooting. The learned frame-level attention scores are re-used, without additional supervision, to drive a diversity-aware keyframe selection module that automatically compresses hours of footage into a compact, timestamped evidentiary summary whenever an anomalous segment is detected. The proposed system was evaluated on the UCF-Crime and DCSASS benchmark datasets for detection and on TVSum/SumMe-style protocols for summarization quality. Experimental results indicate that the framework attains a detection accuracy of 98.4%, an F1-score of 0.978, and an AUC of 0.991, outperforming C3D, two-stream CNN, and CNN-LSTM baselines by 3.4-16.0 percentage points, while the attached summarization module achieves an F-score of 0.71 with a video compression ratio exceeding 92% and end-to-end inference throughput of 41 frames per second on a single consumer-grade GPU. These results demonstrate that combining anomaly-aware attention with keyframe extraction in a single trainable pipeline yields both higher detection fidelity and immediately actionable, human-reviewable video summaries, making the framework well suited for real-time deployment in smart-city and public-safety surveillance infrastructures.

Read PDF

Similar papers

#edge computing Open access Sep 2026

Deep Learning-Based Forensic Detection of Suspicious Activities in CCTV Systems

A deep learning-based forensic framework for real-time detection of suspicious human activity in CCTV videos, trained without relying on any external sensors is proposed, and incorporates anonymization of personal identities and local edge-based processing to prevent raw data exposure.

Qazi Mazhar Ul Haq, Muhammad Imran, M. Waqas et al. · 0 citations
Open access Jul 2026

A CNN-transformer multimodal architecture for weakly-supervised audio-visual violence detection

Automated detection of violent events in surveillance-scale video is important for timely intervention, but frame-accurate labels are too costly to collect at scale. This has made weakly supervised learning from video-level labels the dominant paradigm. Most existing methods score short video snippets from a single mod...

Rahul Gyawali, Rabin Dhakal, Deepak Dulal et al. · 0 citations
Aug 2026

CAT-Net: a coordinate attention transformer network for workplace activity recognition

A hybrid deep learning (DL) model that integrates Coordinate Attention (CA), Convolutional Neural Networks (CNN), and Transformer encoders for better HAR achieves superior performance and proved its effectiveness for workplace safety monitoring applications.

E. Jyotsna, T. Jarin · 0 citations
Review Aug 2026

Language-driven models for video anomaly detection: taxonomy, datasets, and future horizons

This work proposes a novel taxonomy categorizing 34 existing language-driven VAD methods based on learning paradigm, core architecture, adaptation strategy, and functional output, and serves as a foundational resource for the VAD community, advancing the understanding and application of language-driven models in addres...

Olfa Saket, A. Ben Aicha, H. Fathallah · 0 citations
Open access Aug 2026

Multi-level violence recognition via hybrid convolutional-attention and recurrent architectures

Violence recognition is an urgent need for real-time human activity identification in surveillance video streams, which is becoming more important in areas including public safety, law enforcement, and security monitoring. Even while visual understanding has come a long way, finding a balance between accuracy and spe...

Manoj Kumar, Birendra Kumar Verma, Sukhendra Singh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.