Skip to content
Conference

Enhancing Human Activity Recognition using YOLO-based Preprocessing in Deep Learning Frameworks

Jul 2026 · 2026 7th International Conference on Smart Systems and Inventive Technology (ICSSIT) · pp. 1442-1447 · 0 citations · 19 references

Abstract

Human action recognition (HAR) plays a crucial role in safety monitoring, intelligent surveillance systems, and human-computer interaction applications. In this study, we evaluate and compare several deep learning architectures for HAR using the Weizmann dataset, using a YOLO-based preprocessing, including CNN, CNN with attention mechanism, MobileNetV2, and InceptionV3. The proposed YOLO-based preprocessing method was specifically designed to enhance feature extraction efficiency by isolating human subjects from background clutter, thereby reducing noise and improving spatial focus. Experimental results demonstrate that the YOLO-based CNN achieved state-of-the-art performance with an accuracy of 99.6%, significantly outperforming the CNN-Attention model (98.6%), MobileNetV2 (96.1%), and InceptionV3 (93.7%). These findings underscore the importance of robust preprocessing techniques and highlight the superiority of the proposed YOLO-based method in handling complex real-world scenarios.

View source

Similar papers

Conference Jul 2026

A Comprehensive Review of Human Activity Recognition Methods: Trends, Challenges, and Future Directions

Human Activity Recognition (HAR) is a fast-growing research area that focuses on identifying human actions using data collected from sensors and vision-based devices. It plays an important role in applications like health monitoring, smart homes, surveillance, sports analysis, and human-computer interaction. In recent years, several methods have been developed to improve the performance of HAR systems using machine learning, deep learning, and hybrid models. This paper presents a detailed review of different methods used in HAR. The study is divided into three main categories: vision-based methods, sensor-based methods, and hybrid approaches that combine both types. Each method is discussed with examples from recent research, along with their advantages and limitations. A comparison is also provided in the form of a table to highlight the performance and challenges of each approach. Although HAR systems have achieved good results in controlled environments, several challenges still remain. These include poor generalization to new users or unknown environments, difficulty in recognizing complex or overlapping activities, dependence on large datasets, and lack of real-time performance. This paper also discusses these research gaps based on recent findings. The future of HAR depends on building more accurate, reliable, and real-time systems that can adapt to different situations. The paper concludes by suggesting possible directions for future work, such as the development of lightweight models, use of standard datasets, better handling of real-time data, and making models more interpretable.

Satveer Kaur, Navneet Kaur Sandhu, Nitika Goyal · 0 citations
Open access Jul 2026

Human activity recognition using CNN–BiLSTM with attention on hip-mounted wearable sensors

A deep learning–based HAR framework utilizing hip-mounted accelerometer and gyroscope signals from the USC-HAD dataset, which contains readings from healthy participants only, is evaluated, providing a more realistic assessment of subject-independent generalization across unseen individuals.

F. Naveed, Hamza Khan, Zaki Uddin et al. · 0 citations
Review Open access Aug 2026

AI-based vision techniques for human activity recognition in surveillance videos

The article compares the performance of traditional machine learning techniques with recent deep learning architectures such as CNNs, RNNs, TCNs, and Transformers, based on accuracy, computational cost, and suitability for real-world disorderly plotting.

Disha Deotale, Madhushi Verma, P. Suresh et al. · 0 citations
Open access Jul 2026

Video-based human activity recognition analysis using YOLOv26

Human Activity Recognition (HAR) is a rapidly growing research field in computer vision and has various applications, such as intelligent surveillance systems, sports analysis, healthcare, and human-computer interaction. This study aims to analyze the performance of the YOLOv26 Classification model in recognizing video-based human activities using the UCF101 dataset. This study uses six activity classes, namely JumpingJack, Punch, PushUps, Typing, WalkingWithDog, and WritingOnBoard. The research method is carried out through video frame extraction using a uniform sampling technique, the formation of an image classification dataset, YOLOv26 model training, and evaluation at the frame and video levels. Experiments were conducted using 36 configuration combinations consisting of four YOLOv26 model variants (YOLOv26-n, YOLOv26-s, YOLOv26-m, and YOLOv26-l), three variations in the number of frames (8, 16, and 24 frames), and three variations in the number of epochs (50, 100, and 150 epochs). Video evaluation was conducted using a majority voting approach with accuracy, precision, recall, and F1-score metrics. The results showed that all configurations produced video accuracy above 93%. The best configuration was obtained with the YOLOv26-m model with 16 frames and 50 epochs, which achieved video accuracy of 98.26%, precision of 98.15%, recall of 98.37%, and F1-score of 98.23%. Confusion matrix analysis showed that most predictions were on the main diagonal, indicating the model's ability to distinguish human activities with a low error rate. These results prove that YOLOv26 Classification has excellent performance for video-based human activity recognition and has the potential to be applied to various applications based on automatic human activity analysis.

Santi Rahayu, Imam Riadi, Andri Pranolo · 0 citations
Open access Sep 2026

A Lightweight CNN-GRU Model for Human Activity Recognition with Efficient Edge Deployment Using TFLite

Wearable sensor-based human activity recognition (HAR) has become increasingly popular for applications in health monitoring, fitness, and smart living. But the use of deep learning models on edge devices is still challenging due to limited memory and computational power. In this paper, we develop a resource-constrained CNN-GRU hybrid model for HAR on the WISDM dataset. This architecture uses convolutional layers for spatial learning and gated recurrent units (GRU) for sequence learning. For deployment on the edge, the model is quantized to TensorFlow Lite (TFLite) using float16. Our experiments show that the model achieves an accuracy of 94.07%, while the size of the model is substantially smaller and suitable for deployment on edge devices. Importantly, the TFLite model maintains the same accuracy as the original model, ensuring its suitability for real-time deployment. The extensive assessment through confusion matrices, ROC curves and classification metrics confirms the effectiveness of the model across various activities. The proposed approach offers a balance between accuracy and computational efficiency, enabling real-time HAR on edge devices.

Unknown authors · 0 citations
Open access Jul 2026

Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network

A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.

H. Khan, Altaf Hussain · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.