CloudEdgeVLA is introduced, a cloud-edge policy that treats temporal misalignment as a representation-learning problem and offers a practical route to scalable VLA deployment in which cloud models can grow while edge computation remains lightweight and responsive.
Abstract
Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop control must respond locally despite network delay and jitter. Existing hierarchical and asynchronous policies improve throughput, but their slow-path representations can still arrive stale or require explicit scheduling and delay cues. We introduce CloudEdgeVLA, a cloud-edge policy that treats temporal misalignment as a representation-learning problem. A cloud VLA encodes delayed observations into slowly varying task features, while a lightweight edge head combines the latest available cloud feature with current local vision. During training, current and randomly delayed frames are paired with the same current action target in fresh and stale paths. This objective encourages the cloud representation to preserve task-level information while the edge path supplies state-sensitive corrections, driving emergent specialization. Across four LIBERO suites, CloudEdgeVLA retains 63.8-78.0% success with a 40-step uniform-delay window, whereas VLASH reaches at most 6.4% and the evaluated single-path baselines at most 3.0%. By removing blocking synchronization from the control loop, the design offers a practical route to scalable VLA deployment in which cloud models can grow while edge computation remains lightweight and responsive.
A risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested is presented, in which onboard traffic assessment served as a practical trigger for selective VLM inference in these experiments.
Meng Ma, Shu-Yang Li, Nai-Gang Wang et al.· 0 citations
Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...
Rithvik Jonna, Man Namgung, Aakash Gurram et al.· 0 citations
Vision-Language-Action (VLA) models are increasingly offloaded to edge servers, making visual transmission critical for continuous robotic execution. Under packet-loss conditions, delayed or incomplete visual delivery may postpone the generation of subsequent action chunks. However, visual transmission in edge-assisted...
Richeng Huo, Yan-Sen Wang, Ye Wu et al.· 2026 IEEE/CIC International...· 0 citations
This work proposes a framework that exposes backbone depth V, action expert depth A, and denoising steps $D$ as three jointly configurable compute axes in a VLA, and introduces a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deep...
Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al.· 0 citations
SparkVLA is presented, a stop-aware hierarchical VLA that resolves this mutual dependency by formulating both decisions as a single ranking: Stop competes against every action-prefix length in a unified candidate set, and the system selects the highest-scoring option, eliminating threshold tuning and requiring only off...
Xun-Yao Lei, Ren-Jun Wu, Tianlin Huo et al.· 1 citation
Reactive vision-language-action (VLA) policies suffer from task-state aliasing in long-horizon manipulation, where identical multimodal inputs call for distinct, context-dependent actions. Given that pretrained VLAs already possess rich control primitives to express diverse behaviors, we hypothesize that the execution...
Heng-Yan Liu, Wen-Lve Zhou, Bo Yue et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 30, 2026
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.