On-Policy Self-Distillation for Multi-Turn Image Editing
MT-OPSD is proposed, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations, and substantially improves long-horizon editing success and reduces multi-turn collapse.