This paper proposes a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence based on a unified representation via a Decoupled Diffusion Transformer, and closes the fidelity-horizon gap by jointly improving local sharpness, motion and long-range consistency.
Shengqu Cai, Weili Nie, Chao Liu et al.· arXiv.org· 12 citations· ⚡2
Flex-Forcing is introduced, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes, and achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule.
Xinyin Ma, Julius Berner, Chao Liu et al.· arXiv.org· 1 citation
Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achievin...
Neta Shaul, Chao Liu, Arash Vahdat et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.