Skip to content
Preprint

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

Aug 2026 · 0 citations · 56 references
Computer Science

TL;DR

The real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.

Abstract

Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires collecting more data and fine-tuning. We present StellaVLA, a framework that instead adapts at test time by conditioning on a single retrieved demonstration. The key idea is to move beyond imitating what an expert did and instead convey why: an automated offline pipeline converts each raw trajectory into a structured demonstration, e.g., a task plan, sub-goal descriptions, and verbalized 3D motion, at zero human-annotation cost. Provided as in-context guidance, this structured demonstration lets the policy reason about the task rather than mimic a pixel trajectory, which also makes it transferable across embodiments (real-robot, human-hand, or XR demonstrations). A parallel dual-training design internalizes this reasoning during training through a joint action-and-language objective, while inference uses the action expert alone, preserving real-time, high-frequency control with no added latency. On the VLA-Arena leaderboard(Aug 1, 2026), StellaVLA ranks first with an overall score of 0.63, versus 0.44 and 0.22 for the strong prior models ($\pi_{0.5}$ and LingBot-VLA), and it further leads on LIBERO with 98.8% average success rate and LIBERO-Plus with 85.1% success rate. Our real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.

View source

Similar papers

Preprint Aug 2026

How Should Vision-Language-Action Models Use Proprioceptive State?

Five representative interfaces are implemented -- discrete state prompt, VLM prefix, action prefix, state expert, and feature modulation -- under matched implementation details, and evaluated on 45 atomic tasks spanning three task families plus 20 composite tasks.

Yiren Zhao, Ziyang Chen, Ziyang Rao et al. · 0 citations
Jul 2026

DiMaS: Distribution Matching for Steering Vision-Language-Action Models

DiMaS is proposed, a Distribution-Matching Steering strategy tailored to flow-matching VLAs, which transports between representation distributions rather than shifting along a fixed direction, and it effectively controls behavior across two state-of-the-art VLAs.

Pegah Khayatan, Sara Meziane, Jayneel Parekh et al. · 0 citations
Jul 2026

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model

GN (Pangu Navigator), an offline VLN action-prediction system built on OpenPangu-7B, combines mixed-precision computation, selective FP32 computation, and DeepSpeed ZeRO-2 on eight Ascend 910B NPUs and reports metrics quantify offline expert-action alignment rather than closed-loop navigation success.

Li Xian, Mingxi Li, Yizheng Wang et al. · 0 citations
Jul 2026

Semantic Anchoring for Robotic Action Representations

This work examines whether a robot's action representations preserve the semantic structure captured by pretrained encoders and introduces a plug-and-play method that anchors action representations to a semantic manifold while decomposing representations into a shared semantic channel and a private channel, all discarded at inference, leaving the deployed model unchanged.

Yuan Xu, Youheng Shi, Chengyang Li et al. · 0 citations
Preprint Aug 2026

G0.5: One Autoregressive Stream for Robot Reasoning and Action

The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7\% vs.\ 53.3\% for $\pi_{0.5}$ and 24.4\% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4\% vs.\ 26.3\% for $\pi_{0.5}$ and 26.1\% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5\%), a language-following Pick-and-Place benchmark, LIBERO (98.9\%), RoboTwin 2.0 (93.3\%), and SimplerEnv-Bridge (87.3\%).

Yicheng Liu, Zibin Dong, Baijun Ye et al. · 4 citations · ⚡1
Preprint Aug 2026

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

Sangoh Lee, Sangwoo Mo, Wook-Shin Han · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.