Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical...
Zihan Wang, Zhen Wu, Pieter Abbeel et al.· 0 citations
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally...
Kush J. Hari, Justin Kerr, Nidhya Shivakumar et al.· 0 citations
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, pre...
Guan-Ning Zeng, Jia-Ni Wang, Wenjie Ma et al.· 0 citations
This work introduces ABC, a fully open-source stack for manipulation with behavior cloning, and offers a co-training recipe that produces correlated simulation and real-world evaluation, offering a reliable proxy for ablating model-design and training decisions before costly real-world evaluation.
Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.