Skip to content

TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL

Sep 2026 · 0 citations · 37 references
Computer Science

TL;DR

TIDE dynamically rebalanced teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns, and experiments support the effectiveness of TIDE's adaptive OPD--RL coordination.

Abstract

Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.

View source

Similar papers

Preprint Aug 2026

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

This work proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward.

Chen Yang, Hai-Yuan Wan, Rengrong Xiong et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provid...

You-Ling Huang, Tian-Kuo Xu, Jia-Ji Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own generated trajectories. However, teacher supervision can be unreliable when conditioned on incomplete or low-quality student prefixes. We identify...

Jin-Gang Zhou, Yu-Yi Zhou, Hai-Yang Guo et al. · 1 citation
#artificial intelligence Preprint Sep 2026

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d...

Yi-Tong Qiao, Tian-Tian He, Lei Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Reward-Aligned Reweighting for On-Policy Distillation

On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however,...

Hao Xu, Junwei Su, Lan-Song Diao et al. · 0 citations
Preprint Aug 2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al. · 7 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.