Multimodal Contact-Semantic Extraction and Structural Consistency Learning via Embodied Closed-Loop Feedback for Humanoid Robot Assembly
Humanoid robots increasingly rely on multimodal perception and closed-loop interaction to perform contact-rich industrial assembly. However, existing vision-language-action methods often lack fine-grained contact-semantic understanding, explicit structural consistency modeling, and reliable recovery from disturbances....