Skip to content

Multimodal Contact-Semantic Extraction and Structural Consistency Learning via Embodied Closed-Loop Feedback for Humanoid Robot Assembly

Aug 2026 · International Journal of Humanoid Robotics · 0 citations

Abstract

Humanoid robots increasingly rely on multimodal perception and closed-loop interaction to perform contact-rich industrial assembly. However, existing vision-language-action methods often lack fine-grained contact-semantic understanding, explicit structural consistency modeling, and reliable recovery from disturbances. This paper proposes EFSC-Learn, an embodied feedback-driven framework that integrates RGB-D vision, tactile sensing, force-torque signals, proprioception, and human feedback. The method combines temporal alignment, reliability-aware multimodal fusion, contact-semantic extraction, graph-based structural consistency learning, recurrent structural memory, and safety-constrained action projection to generate corrective actions. Experiments on four public robotic manipulation datasets compare EFSC-Learn with ACT, Diffusion Policy, Octo, and OpenVLA. The results show that the proposed method consistently improves contact-semantic recognition, structural relation prediction, assembly success, perturbation recovery, correction accuracy, and contact safety while retaining practical online efficiency under heterogeneous sensing conditions and contact uncertainties. These findings demonstrate that explicit multimodal structure reasoning provides a robust basis for closed-loop humanoid robot manipulation in complex industrial environments.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.