Skip to content

Author

Shengyi Liao

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretr...

Yan-Sen Han, Shengyi Liao, Yuan-Xing Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

This work introduces FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts and shows under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data.

Yan-Sen Han, Shengyi Liao, Peng Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.