Stage-Aware Vision-Language-Action Model for Long-Horizon Robotic Manipulation
Abstract
Vision-Language-Action (VLA) models learn generalist robot manipulation policies by mapping language instructions and visual observations to continuous actions through imitation learning. However, their performance degrades on long-horizon tasks, particularly when sub-tasks admit multiple valid execution orders. Since different demonstrators naturally choose different permutations, the same visual observation is paired with actions for different sub-tasks across the training set. This one-to-many ambiguity introduces contradictory supervision signals, causing the learned policy to oscillate among sub-tasks at inference time. To address this, we propose Stage-Aware VLA, which conditions the policy on an explicit progress state that tracks sub-task completion at runtime, disambiguating the observation-action mapping. Specifically, we introduce a learnable stage query that is jointly processed with visual and language inputs by the VLM to predict the currently active sub-task. Based on this prediction, a stage manager updates the progress state accordingly to guide the policy through the task stages. Experiments on a desk organization task with four sub-tasks show that Stage-Aware VLA improves overall success rate and yields more balanced sub-task completion.