Current music-to-dance generation methods mainly rely on musical features, limiting precise control over generated movements. In particular, most existing methods with control mechanisms do not support example-based control, in which a user provides a reference motion sequence and the generated dances follow its fine-g...
Meng-Qi Liu, Hao Gao, Haolun Li et al.· IEEE Transactions on Visuali...· 0 citations
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatia...
Jia-Ming Tan, Ming-Liang Zhai, Zhen Li et al.· 1 citation
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly rega...
Zhen Li, Zian Meng, Shuwei Shi et al.· arXiv.org· 2 citations
The new version of AlayaWorld substantially revise how conditioning signals are represented and integrated into the model, replacing the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer.
AlayaWorld Team Kaipeng Zhang, Chuan-Hao Li, Yi-Fan Zhan et al.· 5 citations· ⚡2
AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 s...
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.