Skip to content

NAC: Neural Action Codec for Vision-Language-Action Models

Jun 2026 · arXiv.org · Vol abs/2606.21372 · 0 citations · 53 references
Computer Science

TL;DR

The Neural Action Codec is introduced, which treats short robot action trajectories as multi-channel 1D signals and compresses them using a multi-scale RVQGAN architecture and achieves high reconstruction fidelity and higher average success rates than binning, FAST, and prior VQ-based tokenizers at comparable or better compression rates.

Abstract

Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this design through the lens of neural audio codecs - convolutional encoder-decoder architectures with residual vector quantization that serve as the standard front end for audio foundation models. Motivated by their success, we introduce the Neural Action Codec (NAC), which treats short robot action trajectories as multi-channel 1D signals and compresses them using a multi-scale RVQGAN architecture. With adaptations to the action representation, compression rate, and reconstruction objective, audio-codec-style models can autoencode actions with high fidelity without substantial architectural changes. NAC provides a compact, ordered token space via offset codebooks, enabling standard autoregressive policies to operate over short, structured sequences. Meanwhile, a Vocos-style decoder with an ISTFT head and adversarial discriminators recovers action trajectories. Across LIBERO-10, RoboMimic, and a suite of real-world manipulation tasks, NAC achieves high reconstruction fidelity and higher average success rates than binning, FAST, and prior VQ-based tokenizers at comparable or better compression rates. These results demonstrate that repurposed neural audio codecs offer a strong, practical backbone for learned action tokenization in modern VLAs.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.

Shijie Lian, Bin Yu, Zhao-Long Shen et al. · 1 citation
#artificial intelligence Preprint Sep 2026

M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

Experimental results demonstrate the proposed ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance, achieves superior reconstruction fidelity and significantly boosts the success rate of VLA models.

Chun-Pu Xu, Zhi-Xuan Liang, Yu-Hao Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models

Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and n...

Yu-Xin Yang, Gao-Han He, Chang-Xue Guan et al. · 0 citations
#machine learning Preprint Oct 2026

SUAVE: Unified Video-Action Models via Masked Diffusion

Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion...

Rhythm Syed, Jean-Pierre Mercat, Sedrick Scott Keh et al. · 0 citations
#machine learning Preprint Sep 2026

Less Language, More Latents: Annotation-Efficient VLAs for Driving

A three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control and achieves a Driving Score of 87.98 and a Success Rate of 70.46%, matching or surpassing fully supervised baselines on the closed-loop Bench2Drive benchmark.

Alexey Zakharov, Kemal Oksuz, P. Dokania · 0 citations
Preprint Oct 2026

CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation

Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration...

Jin Hyun, Jung Gyu Min, Gyuhyun Jung et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.