Jul 2026· International Conference on Robotics and Sensor Networks· Vol 14254, pp. 1425408 - 1425408-7· 0 citations· 18 references
Engineering
TL;DR
The Adaptive Joint Feature Enhancement (AJFE) module, which learns joint-specific importance weights for improved spatial aggregation, and the Multi-Scale Temporal Sensitivity Enhancement (MTSE) module, which captures multi-scale dynamics via parallel convolutional branches with adaptive fusion are introduced.
Abstract
The Spatial-Temporal Graph Convolutional Network (ST-GCN) has achieved remarkable success in skeleton-based action recognition. However, it exhibits a systematic limitation in fine-grained scenarios, frequently misclassifying subtle, locally driven actions as common whole-body dominant patterns, resulting in persistent semantic confusion. To overcome these shortcomings, we introduce the Adaptive Joint Feature Enhancement (AJFE) module, which learns joint-specific importance weights for improved spatial aggregation, and the Multi-Scale Temporal Sensitivity Enhancement (MTSE) module, which captures multi-scale dynamics via parallel convolutional branches with adaptive fusion. Extensive experiments on NTU RGB+D and Kinetics demonstrate that our approach delivers competitive overall performance (84.5%/88.8% on Cross-Subject/Cross-View protocols) while substantially enhancing semantic consistency in challenging fine-grained cases. These lightweight enhancements offer a practical and effective solution for reliable skeleton-based action recognition in real-world applications.
Skeleton-based action recognition via graph convolutional networks (GCNs) has achieved remarkable progress, yet two persistent bottlenecks limit practical deployment: (1) systematic confusion among fine-grained actions that differ primarily in hand or finger movements, which the standard 25-joint skeleton cannot disamb...
FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.
Imtiaz Ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al.· 0 citations
HPCLR is a hierarchical part-aware contrastive learning framework for skeleton-based action recognition that leverages the consistency among joint, motion, and bone modalities to select more reliable positive samples, thereby contributing to more stable and informative multi-stream skeleton representations.
Hong-Wei Chen, Min Wei, Yuanyuan Zhu et al.· Journal of Supercomputing· 0 citations
This work proposes MEMC (Masked Modeling with Efficient and Minimal Contrastive Learning), a novel framework that adopts an efficient sequential cascade strategy based on layer-grafted pretraining and introduces two CL enhancements to improve the discriminative capability of the learned representations.