Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an...
Hou-Cheng Jiang, Bo-Xuan Zhang, Qi-Yong Zhong et al.· 0 citations
MAS-OPD is presented, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged...
Qi-Yong Zhong, Mao Zheng, Ming-Yang Song et al.· 0 citations
On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-p...
Gengsheng Li, Mao Zheng, Ming-Yang Song et al.· 0 citations
Experiments on reasoning, code-generation, scientific-knowledge, scientific-knowledge, and tool-use benchmarks show that these implementations can be executed through the same verl-based backend while retaining their method-specific objectives and task-dependent performance profiles.
Jie Sun, Mao Zheng, Mingyang Song et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.