Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent...
S. Li, Xiao-Chuang Han, Y. Tsvetkov et al.· 0 citations
The results show that context representations following the principled approach can reduce reliance on model scale, pointing toward a future of AI systems with frontier-level performance powered by smaller models.
Michael Theologitis, Dean Light, S. Li et al.· 0 citations
Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.
Jacqueline He, Howard Yen, S. Li et al.· 0 citations
Translation cascades for reasoning translate the query from another language to English, reason in English, and translate the answer back to the original language. This is a competitive approach to multilingual reasoning, but structurally lossy, since each stage discards information later stages may need, including cue...
Arnav Mazumder, Dengjia Zhang, S. Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.