MoVT is introduced, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation and performs favorably against prior state-of-the-art methods across multiple key metrics.
Abstract
Text-driven 3D human motion generation models face significant challenges in responding to diverse and unconstrained textual prompts, primarily due to the limited availability of 3D motion training data. To address this, we introduce MoVT, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation. At the core of our approach is the cross-modal augmented motion tokenizer, which projects discrete 3D motion tokens into the 2D domain. This projection allows us to enrich the motion codebook with complex, real-world motion patterns derived from videos. The enriched discrete tokens are then mapped back to the 3D domain, resulting in aligned 3D and 2D codebooks with an enhanced capacity to represent intricate motions. These enhanced codebooks are integrated into a generative masked transformer, which predicts masked motion token indices in a modality-agnostic manner. This enables the use of text-index pairs, generated from the 2D codebook and annotated motion videos, to further enhance the generator. Extensive empirical evaluations show that MoVT performs favorably against prior state-of-the-art methods across multiple key metrics.
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotics. Retrieval-Augmented Text-to-Motion (RAG-T2M) improves generation on complex descriptions by conditioning on retrieved motion-text pairs. However, existing RAG-T2M models face two challenges: coarse-grained r...
Yi-Ran Wang, Ze-Yu Zhang, Ling Shao et al.· 0 citations
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as...
Fang-Yuan Tu, Xiang-Yue Zhang, Yi-Yi Cai et al.· 0 citations
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...
Sangeyl Lee, Seunghyun Shin, S. Park et al.· 0 citations
Recent advances in 2D and 3D generative modeling have paved the way for 4D generation. However, existing methods predominantly focus on text-to-4D generation, emphasizing temporal smoothness and spatial consistency while often neglecting precise motion control. Furthermore, due to the inherent ambiguity of text descrip...
Run-Xin Liu, Yang Chen, Ying-Wei Pan et al.· IEEE Transactions on Pattern...· 0 citations
This work investigates a simple yet effective conditional generative framework for text-to-motion generation and proposes T2M-GIT+, which employs a non-autoregressive method to generate discrete motion representations in parallel, and is therefore more efficient than T2M-GPT+ while achieving comparable results.
Jian-Rong Zhang, Yang-Song Zhang, Xiaodong Cun et al.· IEEE Transactions on Pattern...· 0 citations
Text to video generation has advanced significantly in recent years, largely due to the development of extremely sophisticated diffusion models. In this work, we present a novel ap- proach to producing excellent video content based on descriptions by utilizing diffusion tech- niques. Using a multi-stage diffusion proce...
Mohammad Shahnawaz Shaikh· Journal of Intelligent Compu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.