Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while con...
Yuan-Shi Liu, Bo-Yuan Jiang, Liang Hou et al.· 0 citations
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretr...
Yan-Sen Han, Shengyi Liao, Yuan-Xing Zhang et al.· 0 citations
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically eval...
Kai Ding, Xi Chen, Ming-Hong Cai et al.· 3 citations
This work introduces FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts and shows under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data.
Yan-Sen Han, Shengyi Liao, Peng Sun et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.