Preprint
Sep 2026
MoVT: Video-Augmented Motion Tokenizer for Text-to-Motion Generation
MoVT is introduced, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation and performs favorably against prior state-of-the-art methods across multiple key metrics.
Bei-Bei Jing, Tian-Le Guo, You-Jia Zhang et al.
· 0 citations