Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex...
Yu Li, Guang-Feng Cai, Long-Fei Li et al.· 0 citations
Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more...
Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, these methods may produce misleading policy updates because the...
Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al.· 0 citations
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs fr...
Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al.· 4 citations
This paper proposes Disentangled Action--Reasoning Tuning (DART), a simple and efficient framework that explicitly decouples parameter updates for reasoning and tool use via separate low-rank adaptation modules, further supporting the finding of capability interference under shared optimization.
Yu Li, Ming-Yang Yi, Xiu‐Qing Li et al.· arXiv.org· 15 citations· ⚡1
Progress-conditioned Group Policy Optimization is proposed, which uses first-visit observation coverage only when all samples in a group receive zero outcome reward, and consistently improves over group-based baselines, with particularly large gains on hard tasks.
Kaibing Yang, Guangfeng Cai, Sheng-Tian Yang et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.