Jul 2026
Weak-to-Strong Generalization via Direct On-Policy Distillation
Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.
Shiyuan Feng, Huan Gao, Haohan Chi et al.
· arXiv.org · 8 citations
· ⚡1