Diversifying Similar Subjects for Text-to-Image Synthesis with Self-Cross Diffusion Guidance and Reward
This training-free method improves the performance of both U-Net-based and Transformer-based diffusion models, including the Stable Diffusion series and FLUX series and adapts self-cross guidance as an effective reward for RL-based post-training and show improved subject diversity with no computational overhead at inference time.