Scaling Automatic Research Agents via World Models
This paper proposes World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck and accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines.
Xi-Yuan Yang, S. Sarwar, Jingru Cheng et al.
· 0 citations