Skip to content
Preprint

Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

This proposed TDSM-MM has been ablated via extensive experiments and achieved the best inductive accuracy on three of four NTU-60/120 splits and surpasses the transductive state-of-the-art on NTU-120 96/24, suggesting that diffusion-based methods can be a promising direction for zero-shot learning.

Abstract

Zero-shot Skeleton Action Recognition (ZSAR) remains ambiguous when unseen actions share similar skeleton joint dynamics but differ in objects or scene context. RGB provides these missing cues, yet existing multimodal methods typically maintain independent skeleton and RGB scoring branches and fuse their outputs. Without using unlabeled test data for adaptation or fusion calibration, a fixed fusion weight cannot capture class-pair-dependent modality reliability, while an adaptive rule lacks target-side feedback for deciding which branch should dominate. We bypass this weight-selection problem via the classify-by-generation paradigm, where each class is scored by how accurately a text-conditioned denoiser predicts the noise added to the skeleton feature. This formulation separates the progressively corrupted skeleton from fixed conditioning, allowing RGB and text to jointly condition a single class-scoring function rather than produce independent scores. We instantiate this idea as Multimodal Triplet Diffusion for Skeleton-Text Matching (TDSM-MM), augmenting a text-conditioned denoising Transformer with a non-diffused RGB condition token that serves as a stable visual anchor during skeleton data reconstruction. Our proposed TDSM-MM has been ablated via extensive experiments and achieved the best inductive accuracy on three of four NTU-60/120 splits and surpasses the transductive state-of-the-art on NTU-120 96/24 (i.e., 71.3% vs. 69.1%), without test-time adaptation, suggesting that diffusion-based methods can be a promising direction for zero-shot learning.

View source

Similar papers

Aug 2026

Self-supervised skeleton action recognition based on graph prototype learning

This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.

Zhijie Xu, Hong-Wei Chen, Xia Li · 0 citations
Open access Jul 2026

CrossVLS: Cross-Modal Vision-Language Prototypes for Self-Supervised Skeleton Action Representation Learning

Artificial intelligence (AI)-driven positioning and tracking systems combine geometric trajectories with behavior understanding in smart-city, healthcare, and autonomous environments. Skeleton sequences provide compact, privacy-preserving motion geometry, but labeled data are costly, and coordinate-only self-supervision cannot recover object, scene, or interaction cues. Vision-language transfer can supply these cues, but instance-level targets remain sensitive to noisy crops, incomplete descriptions, and ambiguous actions. We propose CrossVLS, which is a cross-modal vision-language-guided framework that transfers semantic knowledge from red, green, and blue (RGB) frames and generated language descriptions to a skeleton encoder during pretraining while retaining skeleton-only inference. CrossVLS replaces noisy instance-level transfer with a shared prototype space: skeleton, RGB, and language features are softly assigned to a common prototype bank through balanced optimal transport, and the resulting assignments define semantic soft targets for contrastive learning. A full-batch progressive training schedule gradually increases cross-modal guidance without splitting the physical batch, preserving the support set used to construct semantic targets. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD demonstrate strong performance under linear and semi-supervised evaluation using only the pretrained skeleton encoder at inference. These results show that prototype-mediated vision-language transfer can improve skeleton representations for the behavior-interpretation stage of positioning and tracking pipelines.

Kenan Ye, Shengjie Zhao, Shuang Liang · 0 citations
Preprint Aug 2026

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.

Imtiaz ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al. · 0 citations
Preprint Aug 2026

GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition

This work proposes GenPrior, the first framework to exploit generative priors from pre-trained Text-to-Motion (T2M) models for ZSAR, and introduces Dispersion-Gated Feature Fusion, which distills kinematic prototypes and intra-class dispersion from generative motion sequences and employs a learned gating network to adaptively inject reliable structural cues into textual embeddings.

Jidong Kuang, Hongsong Wang, Jie Gui · 0 citations
Preprint Aug 2026

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target actions and pretraining using large-scale action scenery datasets. Specifically, our approach, termed Skeleton-Language feature Pooling Switching, introduces a weakly-supervised vision-language pretraining mechanism. This mechanism transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions. Furthermore, we propose Scene-Mixed Discriminative Contrastive Learning to distinguish actions at the instance level within the combined scene through the MIL framework. Our experiments on four public spatio-temporal action localization and classification datasets demonstrate that the proposed method effectively addresses annotation limitations.

Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma et al. · 0 citations
Conference Jul 2026

Cross-modal bidirectional gating for pose-enhanced video action detection

Video action detection requires simultaneous actor localization and action recognition across temporal sequences. Although recent RGB-based methods have achieved strong performance, they often struggle with appearance ambiguity, background clutter, and occlusion. Human pose provides complementary structural cues that are semantically meaningful and less sensitive to irrelevant background information. However, existing RGB-pose fusion strategies are typically oneway, using pose only to guide visual features while ignoring the fact that RGB appearance can also help assess the reliability of pose representations. In this paper, we propose Bidirectional Pose-RGB Modulation Fusion (BPMF), a mutual gating module that enables pose features to guide RGB attention toward action-critical regions while RGB features reweight pose responses to suppress unreliable structural cues. We integrate BPMF into a Mamba-based video action detector with a lightweight pose branch for 2D skeleton modeling. Experiments on JHMDB21 show that our method achieves 76.96% frame-mAP, outperforming both the RGB-only baseline by 4.00 percentage points and the unidirectional fusion variant by 2.84 percentage points. Ablation studies further confirm the effectiveness of the proposed bidirectional design.

Zijun Zhang, Jian-Jun Li, Huangliang Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.