Skip to content

Author

Yao-Hua Tang

We have 3 of 13 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

TCPO: Turn-Level Credit Policy Optimization

Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and iteratively refine their outputs. Although such scores provide dense feedback, they do not directly provide dense credit: a score measures the quality of the current output, while credit should measure how the current turn changes the refinement trajectory. We propose TCPO, a turn-level credit assignment method for verifier-guided multi-turn RL. TCPO casts credit assignment as score-to-credit conversion and constructs turn-level advantages through reference-based comparisons: retrospective credit captures immediate progress and regression relative to the best prior state; hindsight delayed credit identifies non-improving turns with later payoff; and selective fixed-history counterfactual estimation refines high-surprisal turns under the same history. Experiments on math reasoning, code generation, and AppWorld agent tasks show that TCPO improves or matches the strongest baselines across model scales, task domains, and verifier types. TCPO achieves the best or tied-best best-turn Pass@8 on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, reduces turns to success, and improves multi-turn agent performance. These results highlight score-to-credit conversion as a central ingredient for verifier-guided multi-turn policy optimization.

Sicong Liao, Zhi Chen, Yao-Hua Tang · 1 citation
Preprint Aug 2026

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation

This work introduces LEAP (Lean Environment-Feedback via Adaptive Pruning), a scalable and computationally efficient multi-turn RL framework optimized for low-level hardware accelerator alignment and proposes a Rank-Based Reward formulation, establishing a practical paradigm for low-level code RL.

Tankun Li, Zhi Chen, Yao-Hua Tang · 0 citations
Jul 2026

Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

Strong students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union, in an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation.

Songshuo Lu, Zhi Chen, Yao-Hua Tang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.