Jul 2026
EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models
Experiments on reasoning, code-generation, scientific-knowledge, scientific-knowledge, and tool-use benchmarks show that these implementations can be executed through the same verl-based backend while retaining their method-specific objectives and task-dependent performance profiles.
Jie Sun, Mao Zheng, Mingyang Song et al.
· arXiv.org · 0 citations