On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout effici...
Yu-Xiao Yang, Shang-Zhe Li, Tian-Run Yu et al.· 0 citations
We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrep...
Tian-Run Yu, Kai-Xiang Zhao, Shang-Zhe Li et al.· 0 citations
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, th...
Yu-Xiao Yang, Tian-Run Yu, Shang-Zhe Li et al.· 2 citations
TIGER is presented, an inference-time framework that redesigns feedback for localized repair that reduces unsupported content while preserving task quality and a CrisisFACTS case study suggests that the same repair mechanism can improve grounding in multi-source settings.
Kaixiang Zhao, Tianrun Yu, Shawn Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.