Preprint
Aug 2026
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.
Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al.
· 0 citations