Beyond Teleoperation: Enhancing VLA Robustness via Explicit Kinematic Retargeting of Human Demonstrations
Abstract
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalized robotic control, yet their scalability is fundamentally bottlenecked by the high cost and low diversity of teleoperated data. While abundant, human demonstration videos cannot be directly utilized for policy training due to the severe morphological differences between human anatomy and robotic manipulators. To bridge this embodiment gap, this work proposes a lightweight retargeting pipeline that kinematically retargets human interaction data (DexYCB) onto six-degree-of-freedom manipulator trajectories to fine-tune policies based on the pi0.5 architecture. By prioritizing Cartesian positional alignment via constrained Inverse Kinematics (IK) and introducing an object-based grasping heuristic, smooth geometric priors are generated without relying on computationally heavy visual synthesis. Physical evaluations demonstrate that retargeted models significantly outperform standard teleoperation (40.6% success rate), achieving 65.6% success via co-training and a peak 78.1% success rate via two-stage cross-embodiment co-training. Furthermore, evaluations under extreme visual clutter reveal that explicitly retargeted policies exhibit immunity to semantic visual distractors. Finally, we examined and analysed Terminal State Ambiguity, a temporal failure mode where generative models fail to terminate the task when exposed to scenarios similar to the nature of human video priors.