Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

This work proposes RLCSD (Reinforcement Learning with Contrastive on-policy Self-Distillation), which mitigates this drift by contrasting the teacher-student gap under a correct hint against that under a wrong hint, suppressing style shifts induced by hints regardless of correctness and yielding a signal more concentra...

Le-Yi Pan, Shuchang Tao, Yun-Peng Zhai et al. · 30 citations · ⚡4

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.