Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
This work formalizes the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and shows that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate.
Christophe D. Hounwanou, John Emeka Eze, Yaé Ulrich Gaba
· 0 citations