EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
EvoScene-VLA is introduced, which uses compact scene tokens to unify within-chunk scene prediction, cross-chunk state propagation, and observation-based correction and shows that a recurrent scene state alone does not necessarily improve performance, whereas state propagation can improve control when combined with geom...