FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.
Abstract
Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that routes each representation to a content-dependent subset of shared experts, regularized by a load-balancing objective. FineX achieves state-of-the-art results on Gym99, Gym288, and Diving48. On the long-tailed Gym288, it raises mean class accuracy from 68.6% to 76.2% (+7.6 points) without textual supervision or large-scale vision-language pre-training, demonstrating the benefit of structured visual-pose-graph fusion and conditional expert refinement for FHAR.
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hong-Wei Chen, Xia Li· International Journal of Mac...· 0 citations
This proposed TDSM-MM has been ablated via extensive experiments and achieved the best inductive accuracy on three of four NTU-60/120 splits and surpasses the transductive state-of-the-art on NTU-120 96/24, suggesting that diffusion-based methods can be a promising direction for zero-shot learning.
Zehao Bao, Shujun Guo, Bruce X. B. Yu· 0 citations
Cross-domain few-shot facial expression recognition (CF-FER) aims to adapt models trained on basic expressions to recognize novel compound expressions using only a few annotated examples. Although vision-language models (VLMs) have shown promise in few-shot learning, their application to CF-FER remains challenging due to two key issues: coarse-grained textual prompts that fail to capture subtle variations among compound expressions, and episodic training that tends to overfit on highly overlapping few-shot tasks. To address these issues, we propose a fine-grained self-paced relational preserving network (FSR-Net), which introduces fine-grained action unit (AU)-aware textual descriptions generated by large language models (LLMs) to enrich semantic representations and provide more discriminative prototypes. Based on this, we introduce a self-paced relational preserving regularization (SPR) strategy that leverages structural discrepancies between teacher-student visual features and textual-enhanced prototypes as reliability indicators. By progressively weighting reliable samples while filtering out harder ones, the regularization strategy explicitly preserves relational consistency across samples and mitigates overfitting in CF-FER. Comprehensive experiments on multiple CF-FER benchmarks confirm the effectiveness of FSR-Net, yielding average improvements of 5.78% (1-shot) and 3.80% (5-shot) over prior state-of-the-art methods. These results demonstrate its superior capacity for capturing subtle expression cues and enhancing cross-domain transferability.
Kaiyun Wang, Rui Ding, Hanzi Wang et al.· IEEE Transactions on Image P...· 0 citations
Fine-grained video action recognition remains challenging because action categories often differ only in subtle inter-class variations and complex temporal dynamics. Recent Contrastive Language–Image Pre-training (CLIP)-based extensions perform well on general action recognition, but they typically rely on early global pooling of video features. Such coarse representations discard the fine temporal cues that distinguish subtle actions, causing a granularity mismatch in cross-modal alignment. To address this, we propose Structure-Aware Semantic-Adaptive (SASA)-CLIP, a framework for multi-granular cross-modal alignment. SASA-CLIP adopts a dual-branch design: a coarse-grained branch captures the global context, while a fine-grained branch matches descriptors against individual frames before aggregation, rather than pooling features early. To keep this alignment temporally coherent, we introduce a Gaussian prior as a temporal structural constraint, encoding the inductive bias of local temporal continuity into the attention matrix to guide an ordered alignment of key action segments along the temporal axis. On Kinetics-400 (ViT-B/32), SASA-CLIP reaches a Top-1 accuracy of 81.37%, improving over the X-CLIP baseline by 0.97%; on HMDB-51 and UCF-101 (ViT-B/16), it reaches 74.0% and 96.81%, improving by 3.25% and 2.61%, respectively. It also transfers to the zero-shot setting, improving over the baseline on HMDB-51 and UCF-101. These results show that combining multi-granular representations with a temporal structural prior benefits fine-grained recognition, suggesting that SASA-CLIP is a practical option for real-world visual sensing applications such as intelligent surveillance and wearable activity monitoring.
Xiaowei Han, Wenbao Si, Honghui Zhang et al.· Italian National Conference...· 0 citations
Artificial intelligence (AI)-driven positioning and tracking systems combine geometric trajectories with behavior understanding in smart-city, healthcare, and autonomous environments. Skeleton sequences provide compact, privacy-preserving motion geometry, but labeled data are costly, and coordinate-only self-supervision cannot recover object, scene, or interaction cues. Vision-language transfer can supply these cues, but instance-level targets remain sensitive to noisy crops, incomplete descriptions, and ambiguous actions. We propose CrossVLS, which is a cross-modal vision-language-guided framework that transfers semantic knowledge from red, green, and blue (RGB) frames and generated language descriptions to a skeleton encoder during pretraining while retaining skeleton-only inference. CrossVLS replaces noisy instance-level transfer with a shared prototype space: skeleton, RGB, and language features are softly assigned to a common prototype bank through balanced optimal transport, and the resulting assignments define semantic soft targets for contrastive learning. A full-batch progressive training schedule gradually increases cross-modal guidance without splitting the physical batch, preserving the support set used to construct semantic targets. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD demonstrate strong performance under linear and semi-supervised evaluation using only the pretrained skeleton encoder at inference. These results show that prototype-mediated vision-language transfer can improve skeleton representations for the behavior-interpretation stage of positioning and tracking pipelines.
Kenan Ye, Shengjie Zhao, Shuang Liang· Electronics· 0 citations
HPCLR is a hierarchical part-aware contrastive learning framework for skeleton-based action recognition that leverages the consistency among joint, motion, and bone modalities to select more reliable positive samples, thereby contributing to more stable and informative multi-stream skeleton representations.
Hong-Wei Chen, Min Wei, Yuanyuan Zhu et al.· Journal of Supercomputing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.