Skip to content
Preprint

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

A reliable TTT framework for VLA policies (VANE), where candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible.

Abstract

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.

View source

Similar papers

Jul 2026

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

Future-State-Conditioned VLN (FSC-VLN), a deployable model that augments a causal policy with a future-query token and uses training-only future-state supervision to distill information from future observations into the policy state, is proposed.

Lingfeng Zhang, Zhanguang Zhang, Liheng Ma et al. · 0 citations
Preprint Aug 2026

Decoding Task Progress from VLA Representations

The results suggest that VLAs have rich, linearly readable internal representations of semantic quantities like task progress, and that learning to read these signals offers a lightweight, interpretable path toward monitoring deployed visuomotor policies.

Atiksh Bhardwaj, E. W. Duan, Prithwish Dan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations, demonstrating that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies.

Davood Soleymanzadeh, Kai-Di Zhang, Zhiyuan Zhang et al. · 0 citations
Preprint Aug 2026

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation

RA-VLA is presented, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline that facilitates seamless task adaptation while preserving inference efficiency.

Sanghwan Jang, Minjin Jeon, Minsoo Kim et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.