Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
It is found that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones, suggesting that OPD works largely by suppressing low log-probability tokens, which requires no teacher.
Yi Ding, Ruqi Zhang
· 0 citations