Diagnosing On-Policy Self-Distillation for Reasoning Language Models
It is argued that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method because it produces ineffective length growth, stable degradation, or behavioral collapse.
Yang Li, Gong-Le Xue, Yu-Heng Yuan et al.
· 0 citations