A novel contextual bandit algorithm that explicitly incorporates reward decay modeling that achieves significant performance gains over strong baselines and confirms that the integration of reward decay modeling within the bandit framework is crucial for mitigating over-exploitation and optimizing the iterative refinement process.
Abstract
Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect. This neglect leads to over-exploitation, where the continuous use of identical prompts or arms results in diminishing rewards over time. To address this challenge, we propose a novel contextual bandit algorithm that explicitly incorporates reward decay modeling. Utilizing an Expectation-Maximization (EM) algorithm, our method simultaneously estimates both arm-specific and decay parameters. Furthermore, by embedding prompts as arms, we facilitate the joint learning of arm values, distinguishing our approach from the traditional disjoint Linear Upper Confidence Bound (LinUCB) framework. Experimental results on Sentiment Reversal and GSM8K benchmarks demonstrate that our method achieves significant performance gains over strong baselines. Finally, our ablation study confirms that the integration of reward decay modeling within the bandit framework is crucial for mitigating over-exploitation and optimizing the iterative refinement process.
This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.
Xin Shen, Huishuai Zhang, Peng Li et al.· 0 citations
It is shown that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization, and an exact upper bound on the distribution shift induced by a bounded attention update is derived, yielding a trust-region-like analogy to KL-constrained policy optimization.
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. Although promising, these rewards are often noisy and unreliable, limiting their direct utility during deployment. We present Structure-Aware Fine-Tuning (SAFT), a simple, self-supervised method that refines these imperfect reward signals online without access to ground-truth supervision. SAFT leverages intrinsic structural priors to regularize the VLM's latent space via LoRA adapters. We rigorously evaluate SAFT across a spectrum of base model capabilities to demonstrate its versatility. Our results show that SAFT consistently denoises the reward landscape, yielding faster policy convergence and substantially improved alignment (EPIC distance) relative to the underlying base model, suggesting that failures can often be attributed to structural brittleness rather than semantic misunderstanding. By replacing extensive human preference annotation with structural inductive biases inherent to the task, SAFT offers a scalable path for stabilizing text-conditioned RL and underscores the broader value of incorporating task structure as a general inductive bias.
Pyrros Koussios, Chenhao Li, Xin Chen et al.· 0 citations
PAST is proposed, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty and establishes a dual adaptive coordination mechanism that balances the extrinsic and intrinsic rewards.
Ren-Ye Yan, Ji-Kang Cheng, You Wu et al.· 0 citations
High-capacity Click-Through Rate (CTR) models for ads recommendation often exhibit pronounced one-epoch overfitting: performance peaks after a single training pass (epoch) over the data, while additional epochs degrade generalization as the model memorizes noise in high-variance click labels. To address this challenge, we propose TeMPO, a principled framework for generalizable multi-pass training of recommendation models that leverages supervision from a teacher recommendation Foundation Model (FM). Our key idea is to move beyond maximizing the conditional likelihood of noisy click labels. Instead, we maximize the joint likelihood of observing both click labels and the teacher's rich representational knowledge, combining the task loss with teacher-guided alignment objectives. We propose a two-stage optimization strategy to stabilize multi-pass learning. Our theoretical analysis shows that distillation acts as complexity regularization, yielding tighter generalization bounds than standard empirical risk minimization under noisy labels. Extensive experiments on public benchmarks and Meta's production-scale ads application demonstrate that TeMPO enables additional training passes to improve, rather than degrade, performance.
Yunzhe Qi, Qinghai Zhou, Boyang Liu et al.· Proceedings of the 32nd ACM...· 0 citations
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@$k$ behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to interfere. To provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.
Bo-Yang Li, Bingsen Chen, Cheng-Hao Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.