Preprint
Aug 2026
Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
This work proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward.
Chen Yang, Haiyuan Wan, Rengrong Xiong et al.
· 0 citations