This analysis reveals that intermediate log-probability is an unreliable indicator for reasoning capability; instead, reasoning performance results from a shift of internal confidence allocation where RL fine-tuning delays internal convergence, exhibiting prolonged mid-layer exploration before converging sharply at the final output.
Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constr...
Rajat Ghosh, Vaishnavi Bhargava, Henry Wong et al.· 0 citations
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift...
Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised fine-tuning (SFT) can severely overwrite these capabilities. Existin...
Wen-Zhi Fang, Nicholas Tzou, L. Valkov et al.· 0 citations
This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.
Xin Shen, Huishuai Zhang, Peng Li et al.· 1 citation
A systematic literature review on how RL are adapted and scaled as a fundamental post-training tools and how innovations in the RL pipeline enhance the domain-specific LLMs is conducted.
Qianyue Hao, Lin Chen, Xiao-Qian Qi et al.· ACM Computing Surveys· 1 citation
It is found that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.