Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choi...
Hong-Yang Du, Yun-Fei Xie, Jun-Jie Ye et al.· 0 citations
Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their multi-step teachers in generation. However, standard evaluations such as GenEval2 typically draw only one sample per prompt, so improved scores may fail to reveal losses...
Yi-Fei Wang, Xiao-Yu Wu, Tsu-Jui Fu et al.· 0 citations
Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models'own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semanti...
Xiao-Yu Wu, Yi-Fei Wang, Chen Wei· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.