A novel and effective representation is presented that captures action-related information in the pipeline of HAR without any extra annotation overhead beyond the existing skeleton extraction and achieves significantly higher HAR accuracy with similar compactness and efficiency as compared with the state-of-the-art skeleton-only approaches.
Fundamental physical training action recognition plays an important role in intelligent sports analysis, physical education, and motion monitoring. However, RGB-video-based recognition methods are sensitive to background variation, clothing, illumination, and privacy-related constraints, while graph-based skeleton models and pose-heatmap pipelines may introduce additional modeling or preprocessing complexity. To address these issues, this study systematically investigates a lightweight skeleton-rendered pose image pipeline for six fundamental physical training actions. OpenPose was used to extract human keypoints, normalized skeletons were rendered as grayscale pose images, and two-dimensional CNN backbones with temporal aggregation were evaluated under five-fold subject-wise cross-validation. Representative RGB-based, graph-based skeleton, and pose-heatmap baselines were also added for comparison, including R3D-18, Video Swin-T, ST-GCN, CTR-GCN, and PoseC3D. Experimental results show that ResNet18-Max achieved 92.30 ± 0.57% accuracy and 91.88 ± 0.60% macro F1, with 11.69 M parameters, 29.20G FLOPs per 16-frame clip, 18.72 ms inference time, and 53.42 FPS under the implemented setting. The results indicate that skeleton-rendered pose images provide a practical accuracy-efficiency trade-off for structured physical training action recognition.
Shuo Wang, Hao Lu, Ye Zhuang et al.· Discover Applied Sciences· 0 citations
A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.
H. Khan, Altaf Hussain· ICCK Transactions on Advance...· 0 citations
HPCLR is a hierarchical part-aware contrastive learning framework for skeleton-based action recognition that leverages the consistency among joint, motion, and bone modalities to select more reliable positive samples, thereby contributing to more stable and informative multi-stream skeleton representations.
Hong-Wei Chen, Min Wei, Yuanyuan Zhu et al.· Journal of Supercomputing· 0 citations
Human 3D pose estimation is an important problem in computer vision, having applications in a wide variety of fields, such as healthcare, sports science, human-computer interaction, entertainment and retail. However, when solely using 2D sequences of human joints as input data, this becomes an ill-posed problem, caused by the lack of a unique solution, since there are infinitely many 3D points which can be projected to the same 2D point. Current methods rely on temporal or anatomical priors for added information which can help the lifting process. In this work, we propose implementing an additional loss objective for the problem of 3D human pose estimation as a way of embedding semantic information into the lifting process by aligning projected skeleton features with CLIP text embeddings and show that such information can help improve skeleton lifting metrics, especially when applied on out-of-distribution data, without fine-tuning. Our method adds a small MLP projection head to an already existing 3D pose estimation network and a semantic loss between projected features and embeddings of action descriptions, while, at inference, it drops both components, making it adequate for application on novel data, without any prior knowledge of actions. Our code is publicly available at https://github.com/HRIA-HAR/SARPose.
Andrei Mihalea, Mihai Nan, I. Mocanu· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.