Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
AOPD replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning and maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.