Failing to Grasp the Point: Hierarchical Reinforcement Learning for Grasping Tasks
This work studies HIRO-style hierarchy, in which a high-level policy proposes subgoals for a goal-conditioned low-level policy and an off-policy correction relabels past subgoals as the worker improves, and studies an object-centric variant, in which subgoals are defined as relative vectors between task-relevant entities rather than as raw, embodiment-specific robot states.