Skip to content

Author

Zhilong Zhang

We have 1 of 22 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Weak-to-Strong Generalization via Direct On-Policy Distillation

Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.

Shiyuan Feng, Huan Gao, Haohan Chi et al. · 8 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.