Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
Geo-VLA is proposed, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations and introduces Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning.