Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotat...
Jia-Ming Tan, Zhen Li, Shu-Wei Shi et al.· 0 citations
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatia...
Jia-Ming Tan, Ming-Liang Zhai, Zhen Li et al.· 1 citation
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly rega...
Zhen Li, Zian Meng, Shuwei Shi et al.· arXiv.org· 2 citations
The introduction of MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities, is introduced, establishing MM-BrowseComp as a rigorous new standard for the field.
Shilong Li, Xingyuan Bu, Wenjie Wang et al.· arXiv.org· 37 citations· ⚡7
The new version of AlayaWorld substantially revise how conditioning signals are represented and integrated into the model, replacing the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer.
AlayaWorld Team Kaipeng Zhang, Chuan-Hao Li, Yi-Fan Zhan et al.· 5 citations· ⚡2
AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning, and the framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-wit...
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al.· arXiv.org· 1 citation
This work introduces Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems and improves long-horizon consistency and visual quality while retaining real-time streaming throughput.
This work explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance, establishing two properties of Marionette, a world model for interactive games with articulated characters that is directly controllable.
Zian Meng, Zhen Li, Chuanhao Li et al.· 3 citations
AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 s...
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· arXiv.org· 4 citations
Alaya-EVOKE (Evoke) addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation, and achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.