Skip to content
Conference

Adaptive joint weighting and multiscale temporal modeling for fine-grained skeleton action recognition

Jul 2026 · International Conference on Robotics and Sensor Networks · Vol 14254, pp. 1425408 - 1425408-7 · 0 citations · 18 references
Engineering

TL;DR

The Adaptive Joint Feature Enhancement (AJFE) module, which learns joint-specific importance weights for improved spatial aggregation, and the Multi-Scale Temporal Sensitivity Enhancement (MTSE) module, which captures multi-scale dynamics via parallel convolutional branches with adaptive fusion are introduced.

Abstract

The Spatial-Temporal Graph Convolutional Network (ST-GCN) has achieved remarkable success in skeleton-based action recognition. However, it exhibits a systematic limitation in fine-grained scenarios, frequently misclassifying subtle, locally driven actions as common whole-body dominant patterns, resulting in persistent semantic confusion. To overcome these shortcomings, we introduce the Adaptive Joint Feature Enhancement (AJFE) module, which learns joint-specific importance weights for improved spatial aggregation, and the Multi-Scale Temporal Sensitivity Enhancement (MTSE) module, which captures multi-scale dynamics via parallel convolutional branches with adaptive fusion. Extensive experiments on NTU RGB+D and Kinetics demonstrate that our approach delivers competitive overall performance (84.5%/88.8% on Cross-Subject/Cross-View protocols) while substantially enhancing semantic consistency in challenging fine-grained cases. These lightweight enhancements offer a practical and effective solution for reliable skeleton-based action recognition in real-world applications.

View source

Similar papers

Open access Sep 2026

BPA-STGCN: Body-Part-Aware Spatio-Temporal Graph Convolutional Network for Stable Skeleton-Based Action Recognition

Skeleton-based action recognition via graph convolutional networks (GCNs) has achieved remarkable progress, yet two persistent bottlenecks limit practical deployment: (1) systematic confusion among fine-grained actions that differ primarily in hand or finger movements, which the standard 25-joint skeleton cannot disamb...

Xin-Lei Wang, Zhong-Yang Wang, Lu-Xuan Qu et al. · 0 citations
Preprint Aug 2026

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.

Imtiaz Ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al. · 0 citations
Aug 2026

HPCLR: hierarchical part-aware contrastive learning for skeleton-based action recognition

HPCLR is a hierarchical part-aware contrastive learning framework for skeleton-based action recognition that leverages the consistency among joint, motion, and bone modalities to select more reliable positive samples, thereby contributing to more stable and informative multi-stream skeleton representations.

Hong-Wei Chen, Min Wei, Yuanyuan Zhu et al. · 0 citations
Aug 2026

MEMC: Masked modeling with efficient and minimal contrastive learning for self-supervised skeleton-based action recognition.

This work proposes MEMC (Masked Modeling with Efficient and Minimal Contrastive Learning), a novel framework that adopts an efficient sequential cascade strategy based on layer-grafted pretraining and introduces two CL enhancements to improve the discriminative capability of the learned representations.

Yingfei Wu, Wenming Cao, Xinpeng Yin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.