StereoPatch is introduced, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction and suggests that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction.
Ya-Nan Zhou, Zhao-Yan Qian, James Zhao et al.· 0 citations
Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably, leading to spatial misalignment and subsequent manipulation failures. We present MAVP (Map-Aware Visu...
In-hand manipulation allows multi-fingered dexterous hands to reconfigure grasped objects without releasing and regrasping them. This improves manipulation efficiency by reducing repeated grasp acquisition and large arm motions. However, most learning-based methods focus on reorientation, continuous rotation, or transl...
Jun-Xiao Lin, Tian-Yue Wu, Jie Yin et al.· 0 citations
Dexterous manipulation promises substantially richer robot interaction with the physical world, but learning these behaviours remains constrained by the difficulty of collecting consistent, complete-task demonstrations. Unlike parallel-jaw manipulation, dexterous tasks require the operator to coordinate arm motion with...
James Zhao, Jinhe Tang, Mingyuan Ba et al.· 1 citation
Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstration distribution, while the policy continues to produce smooth action c...
This work introduces TRAjectory-routed Causal Evidence (TRACE), a memory framework for visuomotor imitation policies that stores task-relevant visual and robot-state evidence in a fixed-size latent memory that remains bounded over long episodes.
TriManPolicy is presented, a tri-manual imitation learning system that allows one operator to demonstrate behaviours for three arms while reconsidering when they occur, and policies trained on demonstrations retimed by DATS exhibit more efficient coordination while maintaining comparable observed task success.
James Zhao, Mingyuan Ba, Weiming Zhi· arXiv.org· 1 citation
This work proposes a framework known as Video-Generation Environment Representation (VGER), which leverages the advances of large-scale video generation models to generate a moving camera video conditioned on the input image, and demonstrates its ability to produce smooth motions that account for the captured geometry...
Weiming Zhi, Ziyong Ma, Tianyi Zhang et al.· Neural Information Processin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.