Preprint
Aug 2026
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision
This work proposes FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations, and efficiently infers pseudo frame-level alignments without human labeling.
Tongyan Wang, Zhengyuan Li, Muhan Lin et al.
· 0 citations