Jul 2026
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
SEED (SElf-Evolving On-Policy Distillation), a self-evolving framework that converts completed on-policy trajectories into training-time hindsight skills and distills their behavioral effect back into the policy model, is proposed.
Jinyang Wu, Shuo Yang, Zhengxi Lu et al.
· arXiv.org · 7 citations