Vision-Language-Action (VLA) models learn generalist robot manipulation policies by mapping language instructions and visual observations to continuous actions through imitation learning. However, their performance degrades on long-horizon tasks, particularly when sub-tasks admit multiple valid execution orders. Since...
Hao-Ran Shi, Yi-Han Zhou, Ming-Cong Li et al.· 2026 IEEE 22nd International...· 0 citations
This paper presents dynamics-relaxed model predictive control (DR-MPC), a novel MPC formulation for legged locomotion, and a tailored interior-point method (IPM) solver. The formulation combines online optimization feasibility by construction with a contact-aware input parameterization. DR-MPC moves the dynamics equali...
Run Wang, Alapati Tuerxun, Shuo Liu et al.· 0 citations
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide fram...
Yi-Han Zhou, Rui Yan, Ming-Cong Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.