This work proposes a unified visuo-tactile-fusion grasping framework that integrates grasp generation, feasibility prediction, and adaptive refinement and introduces an efficient visuo-tactile representation that tightly fuses object geometry with tactile feedback by associating tactile signals with finger identities.
Abstract
Humans achieve stable and adaptive grasps by seamlessly integrating visual perception and tactile feedback, a capability that remains challenging to replicate in robotic systems. Existing robotic grasping approaches predominantly rely on visual inputs and lack mechanisms for tactile-guided adaptation after contact, limiting robustness and generalization. To address this challenge, we propose a unified visuo-tactile-fusion grasping framework that integrates grasp generation, feasibility prediction, and adaptive refinement. At its core, our method introduces an efficient visuo-tactile representation that tightly fuses object geometry with tactile feedback by associating tactile signals with finger identities. This unified representation supports contact-aware grasp pose generation during planning and tactile-guided refinement after contact, enabling the system to reason about fine-grained finger-object interactions and adjust grasps dynamically. Comprehensive experiments in both simulation and real-world environments demonstrate that our approach significantly enhances grasp success rates and generalization across diverse objects.
This paper presents a tactile-reactive gripper that integrates a Visuo-Tactile Active Palm (VTAP) and compliant, reconfigurable fingers equipped with tactile array sensors. The design exploits structured finger-palm synergy and multi-modal perception to achieve both robust grasping and fine manipulation. The actuated bi-modal palm seamlessly combines long-range visual localization with contact-rich tactile feedback, substantially extending the system's manipulation capability. To bridge the embodiment gap between human hand motion and the heterogeneous three-finger structure, we further propose a staged, gesture-conditioned retargeting framework for dexterous teleoperation. Extensive experiments validate the system across a range of challenging tasks: reactive grasping of YCB and fragile objects, in-hand syringe reorientation and plunger actuation, singulation of clustered objects down to 3 mm in diameter, and vision-tactile peg-in-hole insertion. Results demonstrate that high manipulation performance can be achieved through coordinated finger-palm interaction and multi-modal sensing, without resorting to high degrees of freedom anthropomorphic designs. The VTAP gripper and its retargeting framework offer a practical reference architecture for dexterous gripper design, manipulation, and contact-rich data collection in support of learning-based approaches. Project webpage: https://yuhochau.github.io/vtap/.
Yuhao Zhou, Sheeraz Athar, Zhixian Hu et al.· arXiv.org· 0 citations
The core of the approach is Hybrid Correction Imitation Learning (HCIL), which establishes a “failure-triggered” human-machine mechanism to efficiently resolve the “model gap” via sparse expert corrections.
Changlin Chen, Si-Sheng Chen, Hang Zhang et al.· 0 citations
Recent advances in robotics have highlighted the importance of multimodal perception for dexterous manipulation in contact-rich environments. Here we present BiTAT, a bimanual tactile-augmented teleoperation system for collecting multimodal human demonstrations and learning manipulation policies. The system integrates custom capacitive tactile sensors into parallel grippers and displays the resulting contact-deformation images to the operator. We evaluated the system in four controlled laboratory tasks: USB removal/insertion, bottle cap unscrewing, cucumber peeling, and toothpaste squeezing. In a pilot repeated-measures study with eight laboratory participants, visual tactile feedback was associated with success-rate increases of 12.5–32.5 percentage points and shorter completion times among successful trials. We further propose a multimodal Diffusion Policy that fuses visual, tactile, and proprioceptive features through a Transformer encoder. In two fixed-layout autonomous tasks, the complete model achieved higher observed success rates than the vision-only baseline, including a 45-percentage-point difference in the 50-demonstration toothpaste-squeezing setting. Together, these results demonstrate the feasibility of the proposed hardware–policy pipeline and suggest that tactile augmentation benefits both human teleoperation and learned manipulation policies in contact-rich tasks.
This work proposes a real-world bimanual grasping framework that includes a multimodal dataset capturing joint angles, visual observations and force signals; a Denoising Diffusion Probabilistic Model (DDPM)-based module that generates joint-level grasp configurations from segmented point clouds; and an execution strategy that integrates motion planning with online grasp refinement to ensure physical stability and feasibility.
Ziming Li, Mingxuan Wu, Jiaqi Zhang et al.· 0 citations
Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effectively integrating tactile feedback into dexterous manipulation remains underexplored. In this work, we introduce ReTouch, a vision-language-action model (VLA) that supports contact-rich dexterous manipulation through tactile predictions continually refined online using execution-time feedback. ReTouch builds on two main innovations for tactile representation and closed-loop action generation. First, its Tactile-Patch Encoder represents tactile observations as structured tactile patch features that preserve finger identity and local contact structure, providing contact cues for fine-grained dexterous control. Second, its high-frequency action module jointly predicts future tactile states and action chunks and refines both using incoming tactile feedback during execution. This closed-loop refinement keeps tactile predictions aligned with evolving physical interactions, enabling responsive action correction and improving robustness to contact changes and execution errors. We further introduce XHT-Dataset, comprising 900 real-world demonstrations across seven contact-rich tasks collected on an XHand--UR7e platform, and evaluate ReTouch through closed-loop real-robot experiments. ReTouch surpasses the strongest baseline by 18.4 and 23.8 percentage points in average success rate under standard and challenging conditions, respectively, demonstrating its effectiveness and robustness.
Shiqi Zhang, Xin Zhang, Yedong Shen et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.