Multimodal Contact-Semantic Extraction and Structural Consistency Learning via Embodied Closed-Loop Feedback for Humanoid Robot Assembly
Abstract
Humanoid robots increasingly rely on multimodal perception and closed-loop interaction to perform contact-rich industrial assembly. However, existing vision-language-action methods often lack fine-grained contact-semantic understanding, explicit structural consistency modeling, and reliable recovery from disturbances. This paper proposes EFSC-Learn, an embodied feedback-driven framework that integrates RGB-D vision, tactile sensing, force-torque signals, proprioception, and human feedback. The method combines temporal alignment, reliability-aware multimodal fusion, contact-semantic extraction, graph-based structural consistency learning, recurrent structural memory, and safety-constrained action projection to generate corrective actions. Experiments on four public robotic manipulation datasets compare EFSC-Learn with ACT, Diffusion Policy, Octo, and OpenVLA. The results show that the proposed method consistently improves contact-semantic recognition, structural relation prediction, assembly success, perturbation recovery, correction accuracy, and contact safety while retaining practical online efficiency under heterogeneous sensing conditions and contact uncertainties. These findings demonstrate that explicit multimodal structure reasoning provides a robust basis for closed-loop humanoid robot manipulation in complex industrial environments.