RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
It is demonstrated that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.