Jul 2026· Journal of Science, Innovation and Creativity· Vol 5, pp. 222-241· 0 citations
TL;DR
This research presents a comprehensive review and synthesis of various state-of-the-art deep learning architectures employed in surveillance video-based crime detection and recognition systems, including 3D Convolutional Neural Networks (3D-CNN), Residual Networks (ResNets), Recurrent Neural Networks (RNN), Bidirectional Long- and Short-Term Memory (BiLSTM), Gated Recurrent Units (GRUs), and the integration of attention mechanisms.
Abstract
The rapidly rising crime rates have necessitated advanced and automated security surveillance systems capable of robust, real-time detection, recognition, and prevention of crime. Although traditional surveillance systems have largely been deployed to enhance security and safety, they are inefficient, error-prone, and incapable of effectively processing and generating meaningful insights from the vast quantities of video data they produce. Further, they are adversely affected by extreme weather conditions and subject to human vandalism. The advent of deep learning has significantly transformed earlier automated crime detection, recognition, and prevention by enabling robust extraction and analysis of complex spatial and temporal features from surveillance videos. This research presents a comprehensive review and synthesis of various state-of-the-art deep learning architectures employed in surveillance video-based crime detection and recognition systems, including 3D Convolutional Neural Networks (3D-CNN), Residual Networks (ResNets), Recurrent Neural Networks (RNN), Bidirectional Long- and Short-Term Memory (BiLSTM), Gated Recurrent Units (GRUs), and the integration of attention mechanisms of Soft attention, hard attention, dual attention, and Multi-Head Self-Attention (MHSA). The study critically examines the architectures’ contributions to enhancing detection accuracy, recognition, and the capability to prevent crime. The review further highlights challenges associated with existing systems, including data scarcity, privacy concerns, computational complexity, data class imbalance, and limited real-world adoptability. The paper finally outlines emerging research gaps and future directions in the development of intelligent deep learning-based surveillance crime detection systems.
A deep learning-based forensic framework for real-time detection of suspicious human activity in CCTV videos, trained without relying on any external sensors is proposed, and incorporates anonymization of personal identities and local edge-based processing to prevent raw data exposure.
Qazi Mazhar Ul Haq, Muhammad Imran, M. Waqas et al.· Arab Journal of Forensic Sci...· 0 citations
The paper explores how the classical approaches to image processing have been transformed to deep learning based methods such as their application in object detection and tracking, activity recognition, anomaly detection and facial recognition, and provides the future research direction, which is important to the next-generation intelligent surveillance systems.
Ajay Krishnan· International Journal of Mod...· 0 citations
DeepGuard is presented, an intelligent deep learning framework for automated weapon detection in images and surveillance videos using Faster Region-Based Convolutional Neural Network (Faster R-CNN) and Single Shot Detector (SSD).
Chengoli prashanth, Sk.Mahammadunnisa· American Journal of AI Cyber...· 0 citations
Violence in the public places like fights, accidents, fires and chain snatching has become an alarming situation for the public safety and surveillance systems. Conventional surveillance methods heavily rely on continuous human surveillance which is time-consuming, inefficient, and delayed in response in critical situations. In this paper, we proposed an Intelligent Video Surveillance and Alert System (IVSAS) based on deep learning techniques for real-time violence detection and monitoring to overcome the above limitations. The proposed framework integrates YOLO-based object detection for violent incident identification and MobileNet-based feature extraction for effective spatial feature representation with reduced computational complexity. The system uses OpenCV for live streams of surveillance video and performs keyframe extraction, preprocessing, temporal analysis and confidence-based classification for better detection accuracy with low false positive rates. When violent activity is detected, an automated alert mechanism via the Telegram Bot API immediately sends alert messages and detected incident frames to authorised security personnel for rapid response. The results shows that the proposed system achieves detection accuracy over 90%, real-time processing performance and latency of alert generation. The proposed framework is an efficient, scalable and lightweight solution for real time public safety surveillance applications.
Dedeepya Pulletikurthy, Prasanthi Boyapati· 2026 7th International Confe...· 0 citations
Deep learning has become a key enabling technology for detecting security-relevant events in visual surveillance data acquired from CCTV systems, UAV platforms, and other imaging sensors. However, despite substantial progress in benchmark performance, the operational deployment of such systems remains challenging due to dataset bias, domain shift, limited robustness, edge-computing constraints, and a lack of operationally meaningful evaluation metrics. This structured narrative review synthesises research published between 2021 and 2026, complemented by selected seminal studies. It is guided by predefined research questions and a documented, purposive search and selection strategy, and examines deep learning approaches for weapon detection, violence detection, anomaly recognition, person-related security tasks, perimeter monitoring, and UAV-based surveillance. The analysis covers major architectural paradigms, including YOLO-based detectors, RT-DETR, Vision Transformers, spatiotemporal networks, anomaly-detection frameworks, and edge-optimised models, with particular attention to dataset characteristics, model robustness, adversarial vulnerabilities, multimodal sensing, model compression, and regulatory considerations associated with the EU AI Act. As structuring contributions, the review proposes a multi-dimensional taxonomy that links security-event categories to architecture class, evaluation metric, sensor modality, and EU AI Act risk level, together with a critical analysis of architecture-specific failure modes under operational conditions. It further identifies recurring limitations that hinder real-world deployment: across representative studies, evaluation remains dominated by accuracy-oriented metrics such as mAP, F1-score, and FPS, whereas operational aspects including detection latency, false-alarm burden, and deployment robustness are insufficiently addressed. To bridge this gap, the review proposes Time-to-Detection (TTD) as a complementary operational evaluation framework, and highlights persistent challenges related to realistic datasets, demographic and domain biases, adversarial resilience, privacy-preserving learning, and federated deployment. The findings indicate that future research should prioritise standardised operational benchmarking, TTD-aware evaluation, robust and explainable models, realistic security datasets, efficient edge-AI deployment, and regulation-aware system design. Addressing these challenges will be essential for translating advances in deep learning into reliable and trustworthy security applications.
Martin Havacek, Marek Hutter, Karlis Apalups et al.· Artificial Intelligence Revi...· 0 citations
Security is a significant concern in today’s environment, as unusual activity frequently indicates potential threats and concerns. An abnormality, defined as something that deviates from the expected or normal, can serve as an early warning sign of criminal activity. While crime cannot be anticipated with accuracy, diligent observation of suspect behavior and circumstances can aid in predicting its occurrence. Effective surveillance systems are critical in light of the increasing number of occurrences involving people or groups utilizing firearms to injure or kill. Closed-Circuit Television systems are extensively employed to monitor settings in order to prevent crime, but constant manual surveillance of public locations is difficult. This has resulted in the necessity for intelligent video surveillance, which provides benefits such as effective monitoring, reduced labor, cost efficiency, and the adoption of new security trends. However, human behavior is inherently unpredictable, making it difficult to discriminate between suspicious and typical activity. Early identification of portable weapons such as guns, knives, screw drivers
etc
. are critical for fast response by security officers, potentially lowering violent crimes and homicides. So, in this work a real time human suspicious Monitoring System is proposed using deep learning technique to detect human suspicious activity involving in thefts, crimes and immediately notify law enforcement personnel, thereby improving overall security and safety from the live video stream. This work suggests a real-time method for monitoring suspicious human activity that makes use of the ESP32-CAM and deep learning. The system uses the Haar Cascade model to recognise facial expressions such as sadness, anger and fear. The Yolo model has already been trained to recognise weapons like as knives and guns. The system predicts suspicious conduct such as murder based on the detected weapons and facial expressions. The system predicts theft based on the facial expression. It identifies whether the person is wearing mask and predict the person as suspicious. It notifies law enforcement or authorised people
via
email when a high-probability threat is detected. The system is monitored using an intuitive graphical user interface (GUI), which shows threat statuses, activity logs, and probability-based evaluations for effective management. This method improves on conventional security systems by offering a strong framework for real-time threat identification and response.
V. C., Ajmeera Kiran, H. Alshahrani et al.· PeerJ Computer Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.