Learning to Act under Visual Interruptions with Vision-Language-Action Models
MINT is proposed, which first trains VLA policies to remain functional under missing visual inputs, and selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable.
Ming-Le Jiang, Rui Xu, Yun-Ke Wang et al.
· 0 citations