Skip to content
Conference

Intelligent EfficientNetB7-based Surveillance Video Abnormality Classification Approach with Object Detection and Tracking using YOLOv9 with DeepSORT

Jul 2026 · International Conference Computing Methodologies and Communication · pp. 931-942 · 0 citations · 26 references

Abstract

Video surveillance systems help in tracking and monitoring the real-time events and also anomalous activities. An automated object detection and tracking poses security concerns that minimize the reliance on human intervention. Recent deep learning models provides high productivity and accuracy in dealing with videos of different qualities, however, as surveillance videos have less resolution and poor visibility, more rigorous strategies are required for better tracking of anomaly events. This research introduces object detection and tracking model based on abnormal recognition, enabling better security in crowded areas. Initially, the required videos are gathered through public online databases. The collected videos are directly fed into the object detection and tracking module, where YOLOv9 with DeepSORT (Yv9-DSORT) model is designed to track the objects across the video frames. The detected and tracked frames of the object are passed to the abnormal classification stage, whereas the EfficientNetB7 model is employed to provide abnormality classification results. The classification model identifies complex spatio-temporal patterns and small variations in abnormal regions for threat detection. The developed approach precisely identifies the unusual events. The resultant classified outcomes are validated with the baseline models to ensure its effectiveness.

View source

Similar papers

Open access Jul 2026

Enhanced deep learning model for anomaly object detection and tracking from surveillance videos.

An enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier is proposed using an enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier for anomaly object detection and tracking.

B. Gayal, S. Patil, D. Meshram et al. · 0 citations
Open access Aug 2026

HIGH-PRECISION AERIAL OBJECT DETECTION MODEL UTILIZING YOLO V10 DEEP NEURAL NETWORK

The installation of a real-time visual tracking system with an active pan-tilt camera for indoor human motion detection is presented, which shows that the inclusion of YOLOv10 significantly improves detection precision and temporal consistency.

Ayman Javid Hussain, Lalitha Saroja Ch, Ruqiya Fatima · 0 citations
Open access Aug 2026

Enhancing anomaly detection in video surveillance with spatio-temporal enhanced deep associative memory networks

The proposed STEAD-network combines various techniques, including spatio-temporal enhancement, associative memory modules, and pattern recognition, to effectively capture and recognize abnormal events, and consistently outperforms other methods in anomaly detection accuracy across all datasets.

Kusuma Sriram, Kiran Purusotham, Vinutha Gurulingaiah Kaelgaerae · 0 citations
Jul 2026

Integrated Spatial–Temporal Framework for Video Anomaly Detection in Surveillance Systems

Video-based anomaly detection seeks to discover anomalous events, such as crimes, fires, or medical emergencies, by utilizing both spatial and temporal features of video data. Traditional surveillance systems are frequently limited to minimal recording, requiring human analysts for post-event assessment, resulting in delayed responses during crucial occurrences. To address these issues, we present a multi-layered approach to detecting video anomalies that can deal with both temporal and spatial components of video data. The input video is initially obtained from the dataset and undergoes frame conversion. The extracted key frames are then preprocessed for further analysis. To obtain multi-scale spatial characteristics from each frame, the first layer uses a spatial Pyramid pooling network (SPP-Net) along with a convolutional neural network (CNN). These spatial features are then passed to an optimized bi-directional gated recurrent unit (Opt-Bi-GRU) enhanced with Multi-Head Self-Attention (MHSA), which analyzes the temporal dynamics and captures both forward and backward dependencies across frames. Finally, a capsule network (CapsNet) processes the output of the Bi-GRU, identifying complex patterns that may indicate abnormalities over time. The proposed method is implemented using Python. The proposed model performs better than existing methods in terms of F1-score, specificity, sensitivity, accuracy, recall, precision, FPR, and FNR. The proposed model achieves the highest accuracy of 98.2%, 98.87%, and 98.52%, respectively, utilizing the UBI-fights, UCF-crime, and UCSD pedestrian datasets. These results demonstrate that the proposed framework provides an automated, reliable, and effective solution for real-time anomaly detection in surveillance systems.

M. Rao, Priyesh Kumar · 0 citations
Conference Jul 2026

Query-Driven Intelligent Surveillance using Deep Learning for Activity Recognition and Video Summarization

Surveillance systems have experienced rapid growth which results in production of large video data streams. The monitoring process for this data becomes challenging because its volume exceeds human capacity and this situation creates potential for errors. Our research presents a hybrid intelligent surveillance system which conducts automatic video analysis through its two core operational components. The system employs two primary components to achieve its objectives. The SlowFast-based model enables users to track activities through their development across various time intervals. The system employs YOLO-based models to identify critical objects which include fire and weapons and road accidents through real-time monitoring. The system achieves improved stability through the implementation of a temporal debouncing method. The system uses multiple frame detection checks to improve detection accuracy which helps prevent false alarms. The system includes a module dedicated to video summarization which creates a summary from detected activities and visual changes. The system discards unneeded video content while retaining essential information through this process. The model uses a dataset that contains 4758 video clips which display various classification types. The system reaches 85% validation accuracy which demonstrates its ability to handle new data successfully. The system operates on devices with limited resources while providing an immediate alert system to inform users about essential incidents. The system delivers an easy-to-use and effective solution for intelligent video surveillance operations.

Abdul Haq Nalband, R. U, Shashwat Dodamani et al. · 0 citations
Open access Aug 2026

Object-Centric Industrial Video Anomaly Detection with Local-Global Representation Learning

A novel framework centered on object-centric video anomaly detection, heavily augmented by a local-global representation learning mechanism, suggesting that integrating structured object interactions into representation learning provides a highly scalable and robust solution for real-world industrial monitoring.

C. So, Man-Kit Chau · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.