We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...
Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approac...
Chong-Bo Zhao, Jiangming Wang, Xilai Wang et al.· 0 citations
PILOT (Physical Inference for Latent Optimized Trajectories) bridges the gap between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability of world evolution modeling for action generation.
Xiangkai Ma, Yue Ma, Junjie Wang et al.· 0 citations
LiveLight is presented, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control, which achieves state-of-the-art relighting quality while running at real-time speed, and will publicly release its models, training data, and synthetic data generator.
Yue Ma, Jiangming Wang, Yucheng Wang et al.· 1 citation
Method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data, is presented, demonstrating that the synthesized data substantially improve downstream WAM generalization.
Zexuan Yan, Yuzhou Wu, Yue Ma et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.