A hybrid deep learning (DL) model that integrates Coordinate Attention (CA), Convolutional Neural Networks (CNN), and Transformer encoders for better HAR achieves superior performance and proved its effectiveness for workplace safety monitoring applications.
Experimental results demonstrate that Z-score standardization improves classification performance, and the feasibility and robustness of the proposed framework in real-world traffic environments are indicated.
Dhartee Patel, Jinal Ahir, Namrata Shroff et al.· ITEGAM- Journal of Engineeri...· 0 citations
Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning wit...
CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al.· IEEE Internet of Things Jour...· 0 citations
Intelligent crowd behaviour analysis is critical for modern video surveillance systems to enhance public safety and enable timely anomaly detection. This study presents ViTBN, a vision transformer-based framework augmented with batch normalization to improve feature robustness and the stability of training. The propose...
Ayushi Tiwari, A. S. Kushwaha· International Conference Com...· 0 citations
A deep learning-based forensic framework for real-time detection of suspicious human activity in CCTV videos, trained without relying on any external sensors is proposed, and incorporates anonymization of personal identities and local edge-based processing to prevent raw data exposure.
Qazi Mazhar Ul Haq, Muhammad Imran, M. Waqas et al.· Arab Journal of Forensic Sci...· 0 citations