Mode Seeking meets Mean Seeking for Fast Long Video Generation
This paper proposes a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence based on a unified representation via a Decoupled Diffusion Transformer, and closes the fidelity-horizon gap by jointly improving local sharpness, motion and long-range consistency.