Aug 2026· International Journal of Humanoid Robotics· 0 citations
TL;DR
Experiments show that VTNSCorr achieves the best assembly success rate, rule consistency score, correction success rate, and safety violation reduction, and the results demonstrate its effectiveness for reliable and interpretable industrial humanoid precision assembly.
Abstract
Industrial humanoid robots require robust visual-tactile embodied intelligence to perform precision assembly under occlusion, contact uncertainty, and human-robot coexistence. Existing visuomotor policies and symbolic planning methods often fail to jointly handle continuous contact deviation and discrete industrial rule constraints. This paper proposes VTNS-Corr, a visual-tactile embodied perception-driven neuro-symbolic framework for rule consistency modeling and contact-deviation automatic correction. The method fuses stereo vision, tactile maps, force-torque feedback, proprioception, and human-safety cues into an embodied latent state, grounds it into differentiable symbolic predicates, evaluates fuzzy rule violations, attributes conflict sources, and generates corrective actions such as re-localization, grip adjustment, retreat-and-reinsert, and safety pause. Experiments on RLBench, ManiSkill2, RoboMimic, and DROID show that VTNSCorr achieves the best assembly success rate, rule consistency score, correction success rate, and safety violation reduction. The results demonstrate its effectiveness for reliable and interpretable industrial humanoid precision assembly.
Humanoid robots increasingly rely on multimodal perception and closed-loop interaction to perform contact-rich industrial assembly. However, existing vision-language-action methods often lack fine-grained contact-semantic understanding, explicit structural consistency modeling, and reliable recovery from disturbances....
Hai-Feng Ma, Hui Pan, Lin Li et al.· International Journal of Hum...· 0 citations
The core of the approach is Hybrid Correction Imitation Learning (HCIL), which establishes a “failure-triggered” human-machine mechanism to efficiently resolve the “model gap” via sparse expert corrections.
Chang-Lin Chen, Si-Sheng Chen, Hang Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
In unstructured environments, endowing robots with the ability to dexterously and safely grasp unknown objects presents a critical challenge. Existing control methods struggle to adapt dynamically like human hands, failing to balance grasping stability and object safety. Inspired by human grasping mechanisms, we propos...
Yu-Yao Qi, Tian-Le Wang, Yi-Da Fang et al.· IEEE Robotics and Automation...· 0 citations
Grasping is a fundamental robotic capability that bridges perception and physical task execution. This paper studies grasp pose recovery under a perception-to-execution mismatch, where a grasp generated from visual perception may become spatially stale if the object moves before execution, using only sparse tactile int...
Hao-Ran Wang, Yu-Teng Sun, Yuan-Jie Li et al.· 0 citations
Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition system that estimates object mass and friction-coefficient classe...
Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provid...
Jiaying Chen, Wen-Long Dong, Yan Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.