QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs.
Wei-Qi Wang, Yu-Xin Zhou, Mou-Xiang Chen et al.
· 0 citations