Skip to content
Conference

MMIRL: a multimodal framework for learning robotic manipulation from RGB-D demonstrations

Jul 2026 · International Conference on Hydromechatronics and Advanced Robot Control Technology · Vol 14253, pp. 142530R - 142530R-8 · 0 citations · 6 references
Engineering

Abstract

Extracting robot-reproducible manipulation skills from human RGB-D demonstrations requires stable object identities under multi-object occlusion and a conversion from continuous observations to structured demonstrations. MMIRL addresses this setting with a modular two-stage framework for tabletop manipulation. Offline, the method generates paired synthetic data from a Unified Robot Description Format (URDF) object library and combines it with mixed synthetic data rendered on real backgrounds to calibrate segmentation prompts, candidate selection, and cross-frame association. Online, a detector provides box prompts, a promptable segmentation model predicts instance masks, and identity maintenance performs one-to-one association under joint appearance and geometric constraints to recover identity-aware instance trajectories and 2.5D states. Hand keypoint trajectories then localize interaction intervals and manipulated objects, enabling object-relative normalized snippets for conditional skill and trajectory learning. Evaluations on paired synthetic data and real demonstrations show stable cross-frame association under varying object counts and occlusion, and yield reliable structured demonstrations for replay and execution validation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.