Learning More from Less: Reinforcement Learning from Hindsight
Learning from Hindsight is presented, which brings hindsight relabeling to RL post-training of VLAs by scoring failed rollouts against the tasks they actually achieved, and achieves 5 times improvement in sample efficiency, and outperforms a dense progress-reward baseline on out-of-distribution LIBERO-PRO tasks.