Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
Multitask discriminator Proximity-Guided IRL (MPG) is introduced, which learns two complementary reward components: a generalizable discriminator that transfers shared structure across related tasks to identify expert behavior in a new task and a proximity function that measures how far a state deviates from expert behavior and provides corrective guidance during exploration.