When Do Larger Batches Help Scale LLM Reinforcement Learning?
A larger-batch configuration reduces time-to-target only when its throughput gain exceeds its samples-to-target penalty, and a larger-batch configuration reduces time-to-target only when its throughput gain exceeds its samples-to-target penalty.
Ziniu Li, Jinbo Wang, Guan-Hua Huang et al.
· 0 citations