Preprint
Aug 2026
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
The real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.
Siyu Xu, Yunke Wang, Zijian Wang et al.
· 0 citations