Boosting LLM Exploration via Weak-Model Guidance in RLVR
This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.
Xin Shen, Hui-Shuai Zhang, Peng Li et al.
· 0 citations