On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The ref...
Rui Li, Li-Yang He, Zheng Zhang et al.· 0 citations
CCubicSplat is a differentiable vector rasterizer that replaces B\'ezier closest-point solvers with uniform polyline surrogates whose geometric error is bounded at $O(S^{-2})$.
Chenglong Liu, Xin Zhang, Yi-Meng Zhu et al.· 0 citations
AD-Reranker is proposed, a novel framework that shifts reranker training from proxy imitation to answer-driven utility optimization, and reformulate the reranker as an environment-grounded agent that interacts with a downstream reader, modeled as a deterministic environment.
Keyu Zhu, Shuanghong Shen, Xianquan Wang et al.· Annual International ACM SIG...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.