TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL
TIDE dynamically rebalanced teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns, and experiments support the effectiveness of TIDE's adaptive OPD--RL coordination.
Yi-Bin Huang, Xin-Ming Xu, Cong-Hui Zhu
· 0 citations