#artificial intelligence
May 2026
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
AOPD replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning and maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.
Nan Jia, Haojin Yang, Xing-Chen Ma et al.
· arXiv.org · 18 citations
· ⚡5