Skip to content

Author

Ming-Yang Song

We have 5 of 38 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

MAS-OPD: On-Policy Distillation for Multi-agent Systems

MAS-OPD is presented, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged...

Qi-Yong Zhong, Mao Zheng, Ming-Yang Song et al. · 0 citations
#machine learning Preprint Sep 2026

RAZOR: Pruning Replaceable Experts in LLMs

Mixture-of-experts (MoE) models activate only a few experts per token but store the entire expert pool. Pruning this pool requires identifying experts whose removal preserves model behavior. Routing frequency and output magnitude do not fully describe deletion damage, which also depends on how the surviving and replace...

Ming-Yang Song, Mao Zheng · 0 citations
#machine learning Preprint Sep 2026

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot...

Jie Sun, Mao Zheng, Ming-Yang Song et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Data-free On-policy Distillation

On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-p...

Gengsheng Li, Mao Zheng, Ming-Yang Song et al. · 0 citations
Jul 2026

EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models

Experiments on reasoning, code-generation, scientific-knowledge, scientific-knowledge, and tool-use benchmarks show that these implementations can be executed through the same verl-based backend while retaining their method-specific objectives and task-dependent performance profiles.

Jie Sun, Mao Zheng, Mingyang Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.