Skip to content
Preprint

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos.

Abstract

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.

View source

Similar papers

Preprint Aug 2026

GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics with...

Yanping Zhao, Hang Yu, Yi-Wei Wang et al. · 2 citations
Preprint Sep 2026

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of V...

Xiang-Yu Zhu, Jin Xu, Yue (Sophie) Guo et al. · 0 citations
Preprint Aug 2026

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-r...

Xiaomi Embodied Intelligence Team, University of Macau Shaoqing Xu, Fang Li et al. · 0 citations
Preprint Aug 2026

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-d...

Sen-Qiao Yang, Cheng-Yao Wang, Yuxin Chen et al. · 5 citations
Preprint Aug 2026

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

LAWM-3D is proposed, which introduces three tightly coupled key designs: a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions, a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, and a non-injective RGB-D joint recons...

Jia-Rui Yang, Jiale Zhange, Jiawei Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.