Skip to content
Preprint

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

Persistent Consistency Self-Distillation (PCSD) is proposed, which derives token-level distillation weights from the local persistence of teacher-favoring signals, and combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support.

Abstract

Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated token-level discrepancies, which can be sensitive to noise, or assign a shared step-level weight that may overlook positional variation. We propose Persistent Consistency Self-Distillation (PCSD), which derives token-level distillation weights from the local persistence of teacher-favoring signals. PCSD combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support, applies trend-aware modulation to attenuate locally declining support, and produces continuous weights through sigmoid gating. The resulting objective is jointly optimized with GRPO, combining dense teacher guidance with sparse environmental feedback. Without inference-time skills, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones, exceeding GRPO by 15.6 and 13.3 points and SDAR by 6.2 and 5.5 points, while remaining competitive on WebShop and gaining 15.8 points over GRPO on unseen ALFWorld split.

View source

Similar papers

Jul 2026

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

Group-Reflective Self-Distillation (GRSD), which derives capability-aligned and outcome-discriminative guidance from the policy's own verified rollouts, and refines turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-determined learning direction.

Binbin Zheng, Zi-Jun Xie, Guan-Qun Zhao et al. · 1 citation
Preprint Aug 2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al. · 1 citation
Preprint Aug 2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

Observation-Calibrated Self-Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation-Ablated, to derive an observation residual that discounts score changes shared by the replay scaffold, and applies this residual to modulate token-level GRPO updates at high-uncertainty steps, while preserving the trajectory-level update direction.

Y. Yang, Congming Qin, Xiaodan Liu et al. · 1 citation
Jul 2026

Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

Strong students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union, in an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation.

Songshuo Lu, Zhi Chen, Yao-Hua Tang · 1 citation
Review Aug 2026

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

ADRS is introduced, a framework for constructing return-associated token-level credit for multi-turn language agents that centers and normalizes privileged token scores within each step, modulates them with a return-associated Teacher Value Advantage gate based on within-group confidence--return association, and incorporates the gated token signal into native RL credit construction.

Ran Zhang, Guinan Chen, Chenshaodong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.