On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The ref...
Rui Li, Li-Yang He, Zheng Zhang et al.· 0 citations
CCubicSplat is a differentiable vector rasterizer that replaces B\'ezier closest-point solvers with uniform polyline surrogates whose geometric error is bounded at $O(S^{-2})$.
Chenglong Liu, Xin Zhang, Yi-Meng Zhu et al.· 0 citations
This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.
Cheng Yang, Zhenya Huang, Liyang He et al.· Proceedings of the 32nd ACM...· 0 citations
This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.
Cheng Yang, Zhenya Huang, Liyang He et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.