Skip to content

Author

Ze-Yuan Liu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents

On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used,...

Han-Yang Wang, Ze-Yuan Liu, Zheng-Yu Chen et al. · 0 citations
#machine learning Preprint Sep 2026

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot...

Jie Sun, Mao Zheng, Ming-Yang Song et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Data-free On-policy Distillation

On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-p...

Gengsheng Li, Mao Zheng, Ming-Yang Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.