REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
REOPD combines a token-level compatibility weight with a batch-level adaptive budget, yielding a token-wise coefficient $\lambda_{b,t}=1+\gamma_b q_t$ that preserves teacher alignment while selectively extrapolating along reliable teacher-reference directions, demonstrating effective fine-grained reliability adaptation across domains and teacher configurations.