Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, a...
Xuan Zhang, Long-Tao Zheng, Cun-Xiao Du et al.· 2 citations
It is shown that under GRPO-style optimization, a global normalization baseline may deviate from diverse agents'reward distributions, which ultimately leads to gradient-norm instability, which theoretically pinpoints a key reason for training instability when extending group-based RL to multi-agent LLM systems.
Lang Feng, Long-Tao Zheng, Shuo He et al.· arXiv.org· 14 citations· ⚡1
AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...
Zhiyi Lyu, Ye-Wen Li, Long-Tao Zheng et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.