Skip to content
Open access

Restoring the Right Stream: Training-Free OOD Robustness for Vision–Language–Action Policies

Unknown authors
Sep 2026 · Entropy · 0 citations · 10 references

Abstract

Vision–Language–Action (VLA) policies remain brittle under modest distribution shift. On LIBERO-Plus, contemporary models that solve clean tasks at high rates can fall below 30% success when the camera’s viewpoint or the robot’s initial pose is perturbed. Most training-free test-time remedies address this problem through the image stream, for example, by augmenting, purifying, or selecting visual observations. In our controlled evaluation, this family of methods improves mean success by only about three points and leaves the robot-initial-state failure largely unresolved. This paper studies the failure at the level of input streams. A VLA receives visual tokens, a proprioceptive state token, and language tokens; different perturbations can move different streams away from their training manifold. In particular, the robot-initial-state perturbation directly shifts the proprioceptive token; therefore, image-space interventions have limited leverage. We introduce Gated Per-Stream Manifold Restoration (G-PSMR), a training-free wrapper for a frozen policy. For each stream, a lightweight gate detects off-manifold inputs and applies a stream-specific restoration before the policy forward pass. We instantiate the framework with entropy-gated visual consensus and gated relative-orientation debiasing, which preserves the within-episode orientation trajectory. In the original 280-episode paired evaluation, the joint method improves total success by +5.3 points compared with a +3.2 image-only gain and raises the most fragile factor from 20% to 38%. On 1120 previously unevaluated, manifest-disjoint task instances, the state restoration improves robot-initial-state success from 23.8% to 28.7%; a separate prospectively specified confirmation on 600 new gate-active instances yields 26.3%→32.2% (+5.8 points; 95% CI [+3.2,+8.5]; p<0.001). Together, the original cross-stream results and two independent state-stream evaluations support the central principle of matching the restoration to the input stream carrying the shift.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.