Preprint
Sep 2026
GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments
GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos.
Yi-Chen Liu, Pu-Zhen Yuan, Xiang-Pei Zhu et al.
· 0 citations