Skip to content

Author

Shuo He

We have 2 of 27 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems

It is shown that under GRPO-style optimization, a global normalization baseline may deviate from diverse agents'reward distributions, which ultimately leads to gradient-norm instability, which theoretically pinpoints a key reason for training instability when extending group-based RL to multi-agent LLM systems.

Lang Feng, Long-Tao Zheng, Shuo He et al. · 14 citations · ⚡1
Jul 2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Progress-conditioned Group Policy Optimization is proposed, which uses first-visit observation coverage only when all samples in a group receive zero outcome reward, and consistently improves over group-based baselines, with particularly large gains on hard tasks.

Kaibing Yang, Guangfeng Cai, Sheng-Tian Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.