Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions...
Fu-Kang Liu, Yipu Chen, Jaehwi Jang et al.· 0 citations
Similar to humans, robots benefit from multiple sensing modalities when performing complex manipulation tasks. Current behavior cloning (BC) policies typically fuse learned observation embeddings from multimodal inputs before decoding them into actions. This approach suffers from two key limitations: 1) it requires all...
Vaibhav Saxena, Yun-Hao Luo, Yotto Koga et al.· IEEE Robotics and Automation...· 0 citations
Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video...
Zu-Xing Lu, Hong-Jia Zhai, Guan-Zhi Wang et al.· 0 citations
EgoWAM is introduced, a controlled human-robot co-training framework that fixes the policy backbone, action head, and data mixture while varying only the world prediction target, comparing Pixel, DINO, and 3D motion flow.
Baoyu Li, Xi Yin, Mengying Lin et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.