WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models
WARP-VLA is proposed, a camera-view robust VLA for diverse wrist camera configurations that adopts a Mixture-of-Experts (MoE) architecture where individual experts learn view-specific feature transformations, and a router combines them based on implicit view information.
Junmyeong Lee, Dong-Min Shin, Min-Gyu Park et al.
· 0 citations