Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image gen...
Yuan-Hao Ban, I-Hung Hsu, A. Angelopoulos et al.· 0 citations
Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may only appear in local regions. On-policy self-distillation (OPD) enables transferring privileged visual knowledge from informative views to full-image policies, but queryi...
Zi-Han Chen, Heng-Guang Zhou, Yuan Kang et al.· 0 citations
Improvements are shown that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.
Yi-Hang Chen, Yuan-Hao Ban, Kuei-Chun Kao et al.· 0 citations
An analytical pipeline employing decision trees to discretize continuous neural network attributions into explicit regulatory thresholds is introduced, establishing an auditable methodology to extract robust experimental hypotheses from high-dimensional single-cell data.
J. Chen, Yunqi Hong, Alexandra Bermudez et al.· bioRxiv· 0 citations
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter cre...
Yuanhao Ban, Tong Xie, Sohyun An et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.