The core of the approach is Hybrid Correction Imitation Learning (HCIL), which establishes a “failure-triggered” human-machine mechanism to efficiently resolve the “model gap” via sparse expert corrections.
Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provide rich contact information but introduce additional hardware complexity, calibration requirements, and deployment costs. To bridge this gap, we propose VISTA-Policy, an imitation learning paradigm that utilizes the Visual Deformation Field (VDF), a 3D displacement representation of a compliant gripper, as high-dimensional visuo-physical feedback. The framework integrates: 1) a Physics-Aware Encoding Engine for real-time VDF decoding; 2) an Energy Aggregation Denoising Mechanism to isolate true interaction signals; and 3) a Deformation-Augmented Policy Network with incremental gripper actions for precise closed-loop correction. Extensive evaluations on Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing demonstrate that VISTA-Policy outperforms the strong pure-vision baseline 3D Diffusion Policy and the tactile baseline. VISTA-Policy further demonstrates substantial out-of-distribution generalization to unseen object scales and robustness against dynamic disturbances, offering a durable and cost-effective route toward general-purpose fine-grained manipulation in unstructured environments. Project videos and supplementary materials are available at: https://sites.google.com/view/vista-policy.
Jiaying Chen, Wenlong Dong, Yan Huang et al.· 0 citations
This work proposes a unified visuo-tactile-fusion grasping framework that integrates grasp generation, feasibility prediction, and adaptive refinement and introduces an efficient visuo-tactile representation that tightly fuses object geometry with tactile feedback by associating tactile signals with finger identities.
Xirui Liang, Jiaqi Liang, Jing-Kai Xu et al.· 0 citations
Robust three-finger grasping under physical-domain variation remains challenging because contact stability can change substantially with object mass, effective friction, and observation noise. This work develops U-GRA, a conservative offline-to-online residual adaptation framework for simulated three-finger grasping. U-GRA introduces a unified prior-preserving and critic-disagreement-regulated architecture that couples a frozen behavioral prior with a spectrally normalized and bounded residual stream, scalar Twin-Q reliability assessment, and critic-conditioned residual fusion. The framework first learns a nominal behavioral prior from successful demonstrations and then freezes it as a stable action anchor during online adaptation. Before execution, the twin critics evaluate a candidate action formed from the prior action and the bounded residual proposal, and their absolute scalar Q-value disagreement conditions a state-dependent gate that regulates residual-injection strength. Experiments are conducted in CoppeliaSim using an offline dataset of 40,000 successful demonstrations and online randomization of object mass, effective friction, and observation noise. Across three independent seeds, U-GRA achieves a mean success rate of 84.8±2.3%, a normalized return of 82.7±4.1, and a jitter value of 0.12±0.03. Relative to AWAC-Res, the strongest evaluated baseline, U-GRA improves mean success by 9.2 percentage points and reduces jitter by 57.1%. It also retains the highest mean success rate and normalized return over the unseen simulated high-mass–low-friction OOD region. These results provide simulation evidence that preserving a nominal behavioral prior while regulating bounded residual correction through critic disagreement improves three-finger grasping robustness under physical-domain variation.
Juncheng Zhu, Zhan Gao, Zhile Yang et al.· Machines· 0 citations
In unstructured environments, endowing robots with the ability to dexterously and safely grasp unknown objects presents a critical challenge. Existing control methods struggle to adapt dynamically like human hands, failing to balance grasping stability and object safety. Inspired by human grasping mechanisms, we propose a grasping state regulation strategy based on visual feedforward and tactile gating reflexes to dynamically adjust grasping force. First, guided by the idea that visual information can provide object-dependent expectations before contact, we developed a Two-Stage Mass Estimation Framework and a Vision-Based Friction Coefficient Estimation Framework. They extract the object’s mass and friction coefficients as physical priors from visual inputs, providing reliable initial expectations for subsequent grasping force regulation. Next, the Multimodal State Classifier compares the expected tactile-state representation derived from physical priors and global visual information with the actual tactile-state representation extracted from real-time tactile feedback, and outputs discrete corrective actions. To evaluate performance, we introduce a new metric called the anthropomorphic rate. It quantifies the similarity between the robot-applied force and the human instinctive grasping force. We verified our framework through comprehensive offline and online experiments. Results demonstrate that our system achieves a 91.82% grasp success rate and a 92.65% anthropomorphic rate. These results demonstrate the effectiveness of the proposed strategy in real-world physical interactions.
Yu-Yao Qi, Tian-Le Wang, Yi-Da Fang et al.· IEEE Robotics and Automation...· 0 citations
Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-scale dexterous robot hand data remains difficult. Learning from human demonstrations has emerged as a scalable alternative to robot teleoperation, providing strong priors on object interaction and contact strategies. Recent sim-to-real RL methods incorporate such priors, but often (i) omit rewards that explicitly incentivize precise contact, yielding weak real-world performance, and/or (ii) generalize poorly to unseen object instances. We propose DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points. Its contact-centric rewards encourage precise contact and improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved. Real-world ablations show that DemoMimic achieves 71% success across 16 objects, four tasks, and two robot-hand embodiments, with the smallest sim-to-real drop compared to baselines.
Satvik Sharma, Samrat Sahoo, Huang Huang et al.· 0 citations
Dual-arm robots often encounter difficulties when handling easily deformable or structurally complex objects using traditional grasping-based manipulation. In addition, grasping and releasing operations introduce significant time overhead. To address these limitations, this paper proposes a vision-based predictive control framework for dual-arm nonprehensile transportation. The proposed method employs a hybrid end effector design that integrates an elastic tether with a tray, enabling flexible and stable transportation without direct grasping. A predictive control strategy is adopted to optimize dual-arm motion trajectories on the move under kinematic and safety constraints. To further enhance coordination accuracy, a direct visual servoing scheme is incorporated to dynamically regulate the arm velocities, minimizing relative motion between the end effectors and the object. This effectively suppresses oscillations induced by the elastic tether. Both simulation and experimental results demonstrate that the proposed approach ensures convergence to desired states and achieves continuous, stable, and safe object transportation, even in the presence of disturbances.
Chang Liu, Yuan Yang, Panfeng Huang et al.· 2026 IEEE International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.