Unified Text-Image Generation with Weakness-Targeted Post-Training
This work explores post-training to achieve fully unified text-image generation, where a model autonomously transitions from textual reasoning to visual synthesis within a single inference process, and enables improvements in multimodal image generation across four diverse, independent T2I benchmarks.