Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address this problem. ActNet integrates a Multi-Feature Network (MFNet) backbone for extracting rich features from multiple receptive fields, an Activity Multi-scale Block (AMB) for learning spatially diverse action patterns, and a Focus-Aware Recognition Module (FARM) that adaptively highlights the most informative regions of the image. We evaluate ActNet on the Stanford 40 Actions and PASCAL Visual Object Classes (VOC) 2012 datasets and show that it outperforms several state-of-the-art CNN and transformer-based models, achieving superior accuracy, precision, recall, and F1-score. Extensive ablation studies confirm the effectiveness of both AMB and FARM components. ActNet demonstrates robust generalization to a wide range of human actions, making it a strong candidate for still-image-based action recognition tasks in practical applications.
A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.
Trung-Hieu Le, Thai-Khanh Nguyen, T. Tran et al.· ACM Transactions on Multimed...· 0 citations
Wearable devices play an increasingly pivotal role in human activity recognition (HAR), particularly driven by the urgent demand in medical applications ranging from rehabilitation monitoring to fine-grained gait analysis. However, existing methods still struggle with insufficient exploration of cross-modal information...
Zi-Bo Wang, Runyang Lyu, Bin Xiao· International Conference on...· 0 citations
Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, a...
Xiaowei Han, Ting-Shan Yan, Yunjing Lu et al.· Electronics· 0 citations
Weakly supervised video anomaly detection remains a challenging problem, primarily due to the scarcity of abnormal training samples and the lack of diverse feature representations, which hamper the learning of discriminative models. To address these issues, we introduce a novel weakly supervised cross-domain framework...
Mao-Wen Zhou, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri et al.· PLoS ONE· 0 citations
Action recognition in sports videos remains challenging because of complex motion dynamics, occlusion, and high intra-class variability. Although existing deep learning approaches, including CNN-BiLSTM and transfer learning-based models, have demonstrated effectiveness in human activity recognition, their performance m...
Hai-Ming Yang, Hafiz Mohd Sarim, Xiao-Juan Ma et al.· Journal of Visualized Experi...· 0 citations
Video action recognition requires the joint modeling of spatial appearance information and temporal dynamics. However, existing efficient action recognition methods based on two-dimensional convolution still have limitations in representing multi-scale spatial cues and aggregating key spatiotemporal information. To add...