Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint op...
Zi-Han Su, Junhao Zhuang, Yao-Wei Li et al.· 0 citations
We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation...
Ya-Feng Chen, Bo-Yan Dong, Yan-Kun Huang et al.· 0 citations
Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes...
Zi-Han Su, Si-Wen Lu, Junhao Zhuang et al.· 0 citations
Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...
Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al.· 1 citation
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...
Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin, and the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.
Ding-Yu Yao, Jun Zhou, Chenxu Yang et al.· arXiv.org· 10 citations
A large-scale synthetic data production pipeline built on Unreal Engine for generating action-conditioned, multi-view video, and describes the system architecture, the implementation decisions that emerged from production failures, and the limitations of using perceptual quality proxies for world-model data curation.
Hao-Yu Wang, Song-Chun Zhang, Hao-Ran Li et al.· 0 citations
FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.
Yi-Jun Yang, Shenghe Zheng, Wen-Bo Li et al.· 0 citations
This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.
Xiong-Hao Wu, Yi-Jun Yang, Shi-Long Zhou et al.· 3 citations
The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.
Yi-Cheng Xiao, Wenxun Dai, Xinran Qin et al.· 5 citations
Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory.
Perceptual Flow Matching supervises flow matching in a perceptual feature space using pretrained perceptual models, which substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality.
Chuyang Zhao, Yifei Song, Hongfa Wang et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.