Skip to content

Author

Y. Yang

We have 4 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.

Chishui Chen, Yao Fan, Te Sun et al. · 0 citations
Preprint Aug 2026

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

Influence-Aware Policy Optimization (IAPO), which represents each rollout as a typed influence-dependency graph over trainable agent actions, with user and tool observations serving as evidence, is introduced and advances the understanding of credit assignment in multi-turn user interactions.

B. Ren, Yirong Mao, Y. Yang et al. · 0 citations

Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

SelSkill is proposed, a dual-granularity preference-learning framework for selective skill invocation that formulates skill use as a skill-or-skip decision, uses predictive uncertainty to prioritize candidate decision points, and constructs controlled invoke-skip preference pairs from shared trajectory prefixes.

Chishui Chen, Jiaye Lin, Te Sun et al. · 1 citation · ⚡1
Preprint Aug 2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

Observation-Calibrated Self-Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation-Ablated, to derive an observation residual that discounts score changes shared by the replay scaffold, and applies this residual to modulate token-level GRPO updates at high-uncertainty steps, while preserving the trajectory-level update direction.

Y. Yang, Congming Qin, Xiaodan Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.