Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an...
Hou-Cheng Jiang, Bo-Xuan Zhang, Qi-Yong Zhong et al.· 0 citations
Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot...
Jie Sun, Mao Zheng, Ming-Yang Song et al.· 3 citations
On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-p...
Gengsheng Li, Mao Zheng, Ming-Yang Song et al.· 0 citations
Anchored by this tri-axial framework, representative methods are systematically surveyed, the ongoing transition of continual learning is traced, and the key challenges, broader implications, and future directions arising from this paradigm shift are discussed.
Zhi-Yan Hou, Dan Zhang, Tao Feng et al.· 0 citations
This work presents Language Model Security Modules (LMSM), a security framework that adapts the separation behind Linux Security Modules (LSM) to LLM serving and gives advances in interpretability and model-internal analysis a common path to runtime enforcement.
XiuYu Zhang, Bo-Nan Ruan, Jun-Feng Fang et al.· 0 citations
This work identifies memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance, and proposes AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps.
Mengru Wang, Haozhe Luo, Zhen-Qiang Xu et al.· 0 citations
Mechanist is an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence, and develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretr...
Mengru Wang, Jun-Feng Fang, Shuo-Fei Qiao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.