Tail-Corrected Top-$k$ On-Policy Distillation (TT-OPD) is proposed, which preserves the advantages of TK-OPD, including rich distributional supervision and low computational cost, while providing an unbiased estimator of the gradient of the reverse KL divergence.
Lin-Jian Meng, Si-Yuan Gan, Yu-Hang Li et al.· 0 citations
Direct Advantage Amplification (DAA), which amplifies the advantages of hard-to-sample correct responses on hard prompts, as obtained by Dynamic Sampling, is proposed, which ensures that, when Dynamic Sampling is used, these hard-to-sample responses can be effectively capitalized on, implying higher training efficiency...
Si-Yuan Gan, Yu-Hang Li, Xi-Ran Wang et al.· 0 citations
RA-OPD selects more reliable trajectories to improve student model performance without requiring additional computational cost and is evaluated on math and code benchmarks using models from the Qwen3 family and the DeepSeek-R1 family.
Si-Yuan Gan, Yu-Hang Li, Xi-Ran Wang et al.· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.