Jun 2026· arXiv.org· Vol abs/2606.21372· 0 citations· 53 references
Computer Science
TL;DR
The Neural Action Codec is introduced, which treats short robot action trajectories as multi-channel 1D signals and compresses them using a multi-scale RVQGAN architecture and achieves high reconstruction fidelity and higher average success rates than binning, FAST, and prior VQ-based tokenizers at comparable or better compression rates.
Abstract
Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this design through the lens of neural audio codecs - convolutional encoder-decoder architectures with residual vector quantization that serve as the standard front end for audio foundation models. Motivated by their success, we introduce the Neural Action Codec (NAC), which treats short robot action trajectories as multi-channel 1D signals and compresses them using a multi-scale RVQGAN architecture. With adaptations to the action representation, compression rate, and reconstruction objective, audio-codec-style models can autoencode actions with high fidelity without substantial architectural changes. NAC provides a compact, ordered token space via offset codebooks, enabling standard autoregressive policies to operate over short, structured sequences. Meanwhile, a Vocos-style decoder with an ISTFT head and adversarial discriminators recovers action trajectories. Across LIBERO-10, RoboMimic, and a suite of real-world manipulation tasks, NAC achieves high reconstruction fidelity and higher average success rates than binning, FAST, and prior VQ-based tokenizers at comparable or better compression rates. These results demonstrate that repurposed neural audio codecs offer a strong, practical backbone for learned action tokenization in modern VLAs.
This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.
Shijie Lian, Bin Yu, Zhao-Long Shen et al.· 1 citation
Experimental results demonstrate the proposed ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance, achieves superior reconstruction fidelity and significantly boosts the success rate of VLA models.
Chun-Pu Xu, Zhi-Xuan Liang, Yu-Hao Zhang et al.· 0 citations
Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and n...
Yu-Xin Yang, Gao-Han He, Chang-Xue Guan et al.· 0 citations
Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion...
Rhythm Syed, Jean-Pierre Mercat, Sedrick Scott Keh et al.· 0 citations
A three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control and achieves a Driving Score of 87.98 and a Success Rate of 70.46%, matching or surpassing fully supervised baselines on the closed-loop Bench2Drive benchmark.
Alexey Zakharov, Kemal Oksuz, P. Dokania· 0 citations
Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration...
Jin Hyun, Jung Gyu Min, Gyuhyun Jung et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.