FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel...