Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting...
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for corr...
Wen-Bo Zhang, Peng-Cheng Xu, Wei-Zhi Du et al.· 0 citations
This work proposes Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm that outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time, demonstrating both its effectiveness and efficiency.
Wen-Bo Zhang, Wen-Zhuo Zhou, Heng-Rui Cai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.