Preprint
Aug 2026
XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
This work proposes XCoT-VLA, which replaces descriptive rationales with compact executable CoT tokens learned from automatically constructed Reason-Action supervision, and demonstrates that driving-oriented reasoning can be compact, executable, and directly connected to trajectory generation.
Foundation Model Team, XPeng Inc
· 0 citations