Skip to content

Visual-Tactile Embodied Perception-Driven Neuro-Symbolic Rule Consistency Modeling and Contact-Deviation Automatic Correction for Precision Assembly of Industrial Humanoid Robots

Aug 2026 · International Journal of Humanoid Robotics · 0 citations

TL;DR

Experiments show that VTNSCorr achieves the best assembly success rate, rule consistency score, correction success rate, and safety violation reduction, and the results demonstrate its effectiveness for reliable and interpretable industrial humanoid precision assembly.

Abstract

Industrial humanoid robots require robust visual-tactile embodied intelligence to perform precision assembly under occlusion, contact uncertainty, and human-robot coexistence. Existing visuomotor policies and symbolic planning methods often fail to jointly handle continuous contact deviation and discrete industrial rule constraints. This paper proposes VTNS-Corr, a visual-tactile embodied perception-driven neuro-symbolic framework for rule consistency modeling and contact-deviation automatic correction. The method fuses stereo vision, tactile maps, force-torque feedback, proprioception, and human-safety cues into an embodied latent state, grounds it into differentiable symbolic predicates, evaluates fuzzy rule violations, attributes conflict sources, and generates corrective actions such as re-localization, grip adjustment, retreat-and-reinsert, and safety pause. Experiments on RLBench, ManiSkill2, RoboMimic, and DROID show that VTNSCorr achieves the best assembly success rate, rule consistency score, correction success rate, and safety violation reduction. The results demonstrate its effectiveness for reliable and interpretable industrial humanoid precision assembly.

View source

Similar papers

Aug 2026

Multimodal Contact-Semantic Extraction and Structural Consistency Learning via Embodied Closed-Loop Feedback for Humanoid Robot Assembly

Humanoid robots increasingly rely on multimodal perception and closed-loop interaction to perform contact-rich industrial assembly. However, existing vision-language-action methods often lack fine-grained contact-semantic understanding, explicit structural consistency modeling, and reliable recovery from disturbances....

Hai-Feng Ma, Hui Pan, Lin Li et al. · 0 citations
Conference Open access Sep 2026

PECHC: Robust Tactile Grasping Stabilization in Vision-Denied Peripersonal Space

The core of the approach is Hybrid Correction Imitation Learning (HCIL), which establishes a “failure-triggered” human-machine mechanism to efficiently resolve the “model gap” via sparse expert corrections.

Chang-Lin Chen, Si-Sheng Chen, Hang Zhang et al. · 0 citations
Oct 2026

Human-Inspired Grasping State Regulation Strategy Based on Visual Feedforward and Tactile Gating Reflexes

In unstructured environments, endowing robots with the ability to dexterously and safely grasp unknown objects presents a critical challenge. Existing control methods struggle to adapt dynamically like human hands, failing to balance grasping stability and object safety. Inspired by human grasping mechanisms, we propos...

Yu-Yao Qi, Tian-Le Wang, Yi-Da Fang et al. · 0 citations
Preprint Sep 2026

DA-GRD: Decision-Aware Grasp-Relevant Disambiguation for tactile recovery under perception-to-execution mismatches

Grasping is a fundamental robotic capability that bridges perception and physical task execution. This paper studies grasp pose recovery under a perception-to-execution mismatch, where a grasp generated from visual perception may become spatially stale if the object moves before execution, using only sparse tactile int...

Hao-Ran Wang, Yu-Teng Sun, Yuan-Jie Li et al. · 0 citations
Preprint Aug 2026

ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations

Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition system that estimates object mass and friction-coefficient classe...

Yiwen Liu, Yujun Zhu, K. Jia et al. · 0 citations
Preprint Aug 2026

VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation

Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provid...

Jiaying Chen, Wen-Long Dong, Yan Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.