Skip to content
Preprint

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Sep 2026 · 1 citation · 36 references
Computer Science

TL;DR

Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.

Abstract

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.

View source

Similar papers

Preprint Aug 2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy, which matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.

Jing-Kai Wang, Zihan Tang, Gu Zhang et al. · 0 citations
Preprint Sep 2026

An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM

This work introduces PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch, which allows a single model to predict both how the robot should move and what the scene may look like afterward.

Tian-Heng Wang, Zhou-Yao Xie, Heng-Ji Jia et al. · 0 citations
#machine learning Preprint Sep 2026

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visua...

Yan-Yan Zhang, Di-Sheng Liu, Xin-Peng Li et al. · 0 citations
Preprint Oct 2026

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training e...

Ming-Yu Liu, Chong-Hao Sima, Tianjian Feng et al. · 0 citations
Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

HINT (Human-INTent INcepTion), an agentic framework inspired by the human manipulation principles: semantic intent changes sparsely at manipulation-pattern transitions, whereas continuous control primarily depends on the evolving object-hand relationship.

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.