From Pixels, Without Pre-training: Joint Generative and Self-Supervised Representation Learning in One Model
Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders or autoencoders pretrained for the target domain. We study jo...