GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning
It is found that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.
Outongyi Lv, Yuan-Wei Zhang, Xiao-Qun Zhang
· 1 citation