Skip to content

Author

Guangfeng Cai

We have 3 of 4 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex...

Yu Li, Guang-Feng Cai, Long-Fei Li et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more...

Yu Li, Zheng Zhang, Xin Liu et al. · 0 citations
Jul 2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Progress-conditioned Group Policy Optimization is proposed, which uses first-visit observation coverage only when all samples in a group receive zero outcome reward, and consistently improves over group-based baselines, with particularly large gains on hard tasks.

Kaibing Yang, Guangfeng Cai, Sheng-Tian Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.