Skip to content
Open access

Enhancing anomaly detection in video surveillance with spatio-temporal enhanced deep associative memory networks

Kusuma Sriram Kiran Purusotham Vinutha Gurulingaiah Kaelgaerae
Aug 2026 · Bulletin of Electrical Engineering and Informatics · Vol 15, pp. 3693-3704 · 0 citations · 39 references

TL;DR

The proposed STEAD-network combines various techniques, including spatio-temporal enhancement, associative memory modules, and pattern recognition, to effectively capture and recognize abnormal events, and consistently outperforms other methods in anomaly detection accuracy across all datasets.

Abstract

In the period of extensive video data generation from various surveillance sources, ensuring public safety and security is vital. Detecting unusual crowd behavior is essential, especially in scenarios with large gatherings. However, despite widespread video surveillance, incidents like vehicular accidents, stampedes, and burglaries still occur due to the limitations of traditional surveillance systems. In anomaly detection, finding the critical deviated pattern is the main and critical task. In the context of intelligent video surveillance, automated detection of abnormal behavior is achieved through computer vision analysis, eliminating the need for constant human monitoring. This paper proposes the spatio-temporal enhanced deep associative memory networks (STEAD)-network, a novel approach for anomaly detection in video sequences. The STEAD-network combines various techniques, including spatio-temporal enhancement, associative memory modules, and pattern recognition, to effectively capture and recognize abnormal events. Three benchmark datasets, UCSD Ped2, CUHK Avenue, and ShanghaiTech are used to evaluate this proposed dataset and to compare with existing state-of-the-art techniques. The results demonstrate that the STEAD-network consistently outperforms other methods, achieving significant improvements in anomaly detection accuracy across all datasets. The development of intelligent video surveillance systems is aided by this research by enhancing their ability to autonomously and accurately detect abnormal behavior in real-world scenarios.

Read PDF

Similar papers

Open access Jul 2026

Enhanced deep learning model for anomaly object detection and tracking from surveillance videos.

An enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier is proposed using an enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier for anomaly object detection and tracking.

B. Gayal, S. Patil, D. Meshram et al. · 0 citations
Jul 2026

Integrated Spatial–Temporal Framework for Video Anomaly Detection in Surveillance Systems

Video-based anomaly detection seeks to discover anomalous events, such as crimes, fires, or medical emergencies, by utilizing both spatial and temporal features of video data. Traditional surveillance systems are frequently limited to minimal recording, requiring human analysts for post-event assessment, resulting in delayed responses during crucial occurrences. To address these issues, we present a multi-layered approach to detecting video anomalies that can deal with both temporal and spatial components of video data. The input video is initially obtained from the dataset and undergoes frame conversion. The extracted key frames are then preprocessed for further analysis. To obtain multi-scale spatial characteristics from each frame, the first layer uses a spatial Pyramid pooling network (SPP-Net) along with a convolutional neural network (CNN). These spatial features are then passed to an optimized bi-directional gated recurrent unit (Opt-Bi-GRU) enhanced with Multi-Head Self-Attention (MHSA), which analyzes the temporal dynamics and captures both forward and backward dependencies across frames. Finally, a capsule network (CapsNet) processes the output of the Bi-GRU, identifying complex patterns that may indicate abnormalities over time. The proposed method is implemented using Python. The proposed model performs better than existing methods in terms of F1-score, specificity, sensitivity, accuracy, recall, precision, FPR, and FNR. The proposed model achieves the highest accuracy of 98.2%, 98.87%, and 98.52%, respectively, utilizing the UBI-fights, UCF-crime, and UCSD pedestrian datasets. These results demonstrate that the proposed framework provides an automated, reliable, and effective solution for real-time anomaly detection in surveillance systems.

M. Rao, Priyesh Kumar · 0 citations
Open access 2026

A Deep Spatio-Temporal Framework for Multi-Class Traffic Prediction and Accident Detection in Surveillance Video

Traffic surveillance systems play a crucial role in intelligent transportation by enabling automated monitoring, traffic prediction, and accident detection. However, recognizing complex traffic scenarios from real-world videos remains challenging due to dynamic environments and temporal dependencies. This paper proposes a unified spatio-temporal framework that integrates YOLOv8 based object detection, convolutional neural networks for spatial feature ex traction, and long short-term memory networks for temporal modeling. Traffic videos are preprocessed to enhance visual consistency, and detected objects are transformed into structured spatial representations. Temporal dependencies across video sequences are learned using LSTM networks, and the extracted features are evaluated using multiple machine learning classifiers under different preprocessing strategies. Experimental results demonstrate that Z-score standardization improves classification performance, with Support Vector Ma chine achieving 63.27% accuracy and an F1-score of 55.66% in an eight-class traffic scenario classification task, indicating the feasibility and robustness of the proposed framework in real-world traffic environments.

Dhartee Patel, Jinal Ahir, Namrata Shroff et al. · 0 citations
Open access 2026

FASTe: Framework With Application-Driven Spatio-Temporal Efficiency for Video Anomaly Detection

Video anomaly detection (VAD) plays a crucial role in modern surveillance systems. However, practical deployment remains challenging due to three key limitations: the difficulty of handling variable-length videos, the lack of fine-grained frame-level anomaly localization, and the high computational complexity of using existing models in real-time. To address these challenges, we propose FASTe, a lightweight and spatio-temporal efficient framework for real-time anomaly detection. FASTe introduces 1) a LogSumExp-based multiple instance learning (MIL) aggregation strategy for robust training on variable-length inputs, 2) frame-level anomaly localization under weak supervision, without requiring dense labels, and 3) spatio-temporal decoupling via adaptive pooling, reducing attention complexity from <inline-formula> <tex-math notation="LaTeX">$O(T^{2} \times H^{2} \times W^{2})$ </tex-math></inline-formula> to <inline-formula> <tex-math notation="LaTeX">$O(T^{2})$ </tex-math></inline-formula>. Evaluated on the UCF-Crime dataset, our approach achieves a receiver operating characteristic area under the curve (ROC-AUC) of 94.57%, achieving improved performance under weak supervision. The proposed framework offers a practical and scalable solution for real-time anomaly detection in resource-constrained surveillance environments.

Jihun Jeon, Raeyoung Chang, Jisu Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.