Preprint
Aug 2026
Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts
FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.
Imtiaz ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al.
· 0 citations