PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
PEARL is an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution, and maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time.
Ji-Aan Zhu, Wei Gao, You-Hui Bai et al.
· 0 citations