Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent...
S. Li, Xiao-Chuang Han, Y. Tsvetkov et al.· 0 citations
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study control...
Jia-Xin Ge, Yi-Ming Qin, Ji Xie et al.· 0 citations
This work explores post-training to achieve fully unified text-image generation, where a model autonomously transitions from textual reasoning to visual synthesis within a single inference process, and enables improvements in multimodal image generation across four diverse, independent T2I benchmarks.
Jiahui Chen, Philippe Hansen-Estruch, Xiao-Chuang Han et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.