This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignme...
Tian-Ze Xu, Yan-Zhao Zheng, Zhen-Tao Zhang et al.· 0 citations
CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation, and achieves the highest average among all evaluated student-training methods.
Yan-Zhao Zheng, Yuan-Qiang Yu, Tian-Ze Xu et al.· 0 citations
This work identifies and formalizes the Distribution-Value Coevolution principle: the training value of data is not intrinsic, but emerges dynamically from the interaction between data characteristics and the model's evolving capability boundary, and operationalizes this principle through a unified framework.
Zairun Yang, Yanbo Yang, Chenyi Zhou et al.· Proceedings of the 32nd ACM...· 0 citations
Reinforcement learning from human feedback (RLHF) has become the cornerstone of aligning large language models (LLMs) with human intent. Yet a fundamental question remains unaddressed: how should training data be scheduled when both the model's capabilities and the utility of data are constantly evolving? Current pipel...
Zairun Yang, Yanbo Yang, Chenyi Zhou et al.· Proceedings of the 32nd ACM...· 0 citations
PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance.
Yuan-Qiang Yu, Yan-Zhao Zheng, Zhen-Tao Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.