This work presents vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization, and introduces IMPACT, an ACT-based policy with cached text representations and language-modulated visual features.
Abstract
Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately $1.4\times$ median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield $76.4\%$ mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.
Embodied LLM systems increasingly co-locate latency-critical robotics pipelines with compute- and memory-intensive language-model inference on edge platforms such as NVIDIA Jetson. This co-location avoids cloud round trips and enables privacy-preserving, low-latency interaction, but it also creates a new source of reso...
Teng Mei, Cheng-Xuan Pei, Marco Canini et al.· Proceedings of the 17th ACM...· 1 citation
Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deplo...
LazyTrain is proposed, an optimization layer over a layer-streaming executor that formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training.
Xiao-Jun Wu, Ce-Hao Yang, Hong-Hao Liu et al.· 1 citation
Vision-Language-Action (VLA) models have emerged as foundational models for next-generation robotics. High VLA inference throughput is critical for meeting the control-rate requirements of robots. VLA models comprise two phases, a vision-language model (VLM) and an action head, that can be decoupled and executed asynch...
A. Li, Christina Giannoula, N. Vijaykumar· 0 citations
EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.
Jia-Min Cao, Qingxu Li, Yaozhong Liu et al.· Conference on Applications,...· 0 citations
LLM4LLM is introduced, a deployment-aware closed-loop optimization framework that starts from a target inference script, extracts phase-aware optimization tasks, searches with an experience-guided episodic agent, and accepts patches through in-model validation.
Hui Zeng, Pengfei Yang, Yanxin Chen et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 30, 2026
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.