UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity, supports action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.
Abstract
Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head-wrist observations and paired ego-UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.
JoyAI-RA 0.5 is proposed, a generalist Vision-Language-World-Action framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment, suggesting that abundant but weakly labeled human experience can be converted into a transferable t...
Skel-WAM is introduced, a world action model that bridges differences between human and robot motion through a unified hand-skeleton motion interface, and demonstrates that a shared skeletal interface enables joint learning across human and robot data and expands robot task coverage through complementary human demonstr...
Ze-Tao Cai, Ya-Ping Li, Yi-Qun Wang et al.· 0 citations
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spat...
Song Liu, Lin-Ying Li, Yan-Shun Zhao et al.· 0 citations
EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations, is presented.
Yi-Ming Jiang, Jin Chen, Chong-Yang Xu et al.· 0 citations
Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We pre...
Pengbo Niu, Yu-Jia Xie, Rui Peng et al.· 0 citations
This work evaluates cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated.
Maxime Alvarez, Renzo Caballero, T. Matsushima et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.