Skip to content

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

Jul 2026 · arXiv.org · Vol abs/2607.26094 · 0 citations · 27 references
Computer Science

TL;DR

MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function across auxiliary tasks before RLHF training, is introduced, providing theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization.

Abstract

Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment. We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function $\Phi(x,y;\phi)$ across auxiliary tasks before RLHF training. The learned shaping produces a composite reward that preserves policy optimality while providing task-specific learning signals. Our meta-objective combines task discrimination, entropy regularization, and potential-based conservation for stable convergence. We provide theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization. Experiments on LLaMA-3-8B across four benchmarks show consistent improvements over PPO, DPO, GRPO, and DAPO, achieving a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability. MeRLa retains its benefits when combined with process-based and rubric-based enhanced rewards.

View source

Similar papers

Preprint Aug 2026

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

A unified analytical framework for comparing dynamic reward shaping and neighbouring adaptive reward mechanisms is introduced, which distinguishes parametric revision from state-dependent variation, separates additive shaping from reward replacement and reward-adjacent guidance, and organises existing methods along temporal, informational, and theoretical dimensions.

Fouad Bahrpeyma · 0 citations
Jul 2026

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

This work develops LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies.

Manith Adikari, Bei Peng, Samuele Vinanzi et al. · 0 citations
Jul 2026

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

A principled reward decomposition framework is introduced that decouples verifiability from sparsity, and two new reward functions are proposed: an exponential reward that provides graded penalties based on the count of forbidden-concept occurrences, and a PageRank inspired reward that weights penalties by semantic importance.

Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj et al. · 1 citation
Review

Adaptive Reward Design in Reinforcement Learning: A Taxonomy and Survey

A unified view of ARD in RL is provided by introducing a taxonomy, organized by the primary driver of the reward variation, that distinguishes external-feedback-driven reward updates from reward adaptations driven by endogenous within-run signals and those conditioned on exogenous context signals.

Raphaela Baybas, Carlo D'Eramo, Philipp Brune · 0 citations
Preprint Jul 2026

Cross-Benchmark Generalization in Long-Horizon Agents

Desc descriptive evidence is provided that long-horizon multi-tool post-training can change ways of working that transfer beyond its training domain, and both software-engineering benchmarks improve despite the training collection containing no software-engineering tasks.

Sushant Mehta, Logan Ritchie, Liudas Panavas et al. · 0 citations
Preprint Jul 2026

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents

Reinforcement learning holds significant potential for training large language models to handle multi-turn interactive tasks, but directly training with outcome rewards often results in slow convergence due to the sparsity of signals and the lack of fine-grained feedback.

Qiang Liu, Taian Guo, Ruizhi Qiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.