Jul 2026
Reward-Gated On-Policy Distillation
RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals, and produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline.
Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi et al.
· arXiv.org · 4 citations
· ⚡1