Skip to content
Open access

Temporal Expert-Based Reward Learning for Inverse Reinforcement Learning

Sep 2026 · Jornadas de Automática · 0 citations · 11 references

Abstract

Learning reward functions from expert demonstrations removes the need for manual reward engineering in reinforcement learning applied to robotic manipulation. Existing methods, however, require trajectory quality annotations, episode success labels, or produce implicit rewards that are difficult to inspect. This paper presents TEXB-IRL (Temporal EXpert-Based reward learning for Inverse Reinforcement Learning), which learns a dense neural reward function directly from unlabeled expert demonstrations. TEXB-IRL exploits intra-trajectory temporal ordering through two complementary objectives: a temporal consistency loss that enforces monotonically increasing reward along expert trajectories, and an expert-agent separation loss that anchors the reward scale. The policy is optimized with PPO over the learned reward. Preliminary experiments in simulation with the PAL TIAGo++ robot show competitive results against state of the art algorithms.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.