Manipulation under time constraints requires both accurate actions and an execution rhythm that matches the evolving scene. This becomes critical when a robot must intercept moving objects or complete a sequence of adjustments before a deadline. Although one-step policies reduce generation cost, their directly predicte...
Zhen Dong, Qing-Ran Wu, Jin-Na Fu et al.· 0 citations
A unified four-step post-training pipeline comprising a learned temporal hand-action codec, supervised fine-tuning, DAgger, and real-world residual reinforcement learning provides a practical path for adapting VLA foundation models to reliable real-world dexterous manipulation.
Jun-Lei Zhu, Shen-Zhe Yao, Chao-Gui Huang et al.· 0 citations
This work introduces MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video and develops an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision.
Zi-Jie Zhu, Wei-Ren Cai, Yi-Zhou Wang et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.