Skip to content

Delayed Convergence and Emergent CoT Reliance in RL-Tuned Language Models

· 0 citations · 17 references

TL;DR

This analysis reveals that intermediate log-probability is an unreliable indicator for reasoning capability; instead, reasoning performance results from a shift of internal confidence allocation where RL fine-tuning delays internal convergence, exhibiting prolonged mid-layer exploration before converging sharply at the final output.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

GRPO Training Dynamics for Small Language Models

Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constr...

Rajat Ghosh, Vaishnavi Bhargava, Henry Wong et al. · 0 citations
#artificial intelligence Preprint Oct 2026

On Language Drift during RLVR Post-Training

Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift...

Michael Sullivan, Alexander Koller · 0 citations
#artificial intelligence Preprint Sep 2026

Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised fine-tuning (SFT) can severely overwrite these capabilities. Existin...

Wen-Zhi Fang, Nicholas Tzou, L. Valkov et al. · 0 citations
#small language model Preprint Aug 2026

Boosting LLM Exploration via Weak-Model Guidance in RLVR

This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

Xin Shen, Huishuai Zhang, Peng Li et al. · 1 citation
Review Open access Aug 2026

Reinforcement Learning in the Era of Large Language Models: Challenges and Opportunities

A systematic literature review on how RL are adapted and scaled as a fundamental post-training tools and how innovations in the RL pipeline enhance the domain-specific LLMs is conducted.

Qianyue Hao, Lin Chen, Xiao-Qian Qi et al. · 1 citation
#artificial intelligence Preprint Aug 2026

GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning

It is found that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.

Outongyi Lv, Yuan-Wei Zhang, Xiao-Qun Zhang · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.