Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent...
S. Li, Xiao-Chuang Han, Y. Tsvetkov et al.· 0 citations
The results show that context representations following the principled approach can reduce reliance on model scale, pointing toward a future of AI systems with frontier-level performance powered by smaller models.
Michael Theologitis, Dean Light, S. Li et al.· 0 citations
FLIP (FLipped Inference for Prompt reconstruction), a reference-free and rubric-free reward modeling approach that reformulates reward modeling through backward inference that enables reliable reward modeling in downscaled regimes where judgment methods fail, is proposed.
Translation cascades for reasoning translate the query from another language to English, reason in English, and translate the answer back to the original language. This is a competitive approach to multilingual reasoning, but structurally lossy, since each stage discards information later stages may need, including cue...
Arnav Mazumder, Dengjia Zhang, S. Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.