Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models
Experiments across video understanding and reasoning benchmarks show that the Evidence-Grounded Self-Teacher framework consistently improves upon Standard OPSD across multiple backbones and achieves performance comparable to GRPO while requiring substantially less training time, establishing an effective and efficient post-training approach for Video-LLMs.
Zi-Yue Wang, Shi-Qi Huang, Wei-Wen Xu et al.
· 0 citations