Jul 2026
S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving
The S-squared-VLA is proposed, which explicitly decouples the semantic and spatial streams in Vision-Language-Action models, and significantly outperforms baselines, achieving the highest No Collision rate among all evaluated methods.
Jianguo Yu, Rukang Wang, Duan-Feng Chu et al.
· arXiv.org · 0 citations