Specialize-and-Merge Online Policy Distillation (SMOPD) is proposed, a two-stage training method for multi-reward optimization that outperforms GDPO across 1.5B, 3B and 7B backbones.
Wen Wang, Jiahua Bao, Tu Yongsiqi et al.· 1 citation
NapMem is introduced, a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context, and suggests that long-term user memory benefits from coupling structured storage with a learned policy for using memory at the appropriate granularity.
This work introduces Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, a solver, and a dynamic skill controller that effectively bridges the gap between structured verification and open-ended exploration.
Siyuan Huang, Pengyu Cheng, Haotian Liu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.