Skip to content

Efficient Vision-Language-Action Management and Serving for Robot Factories

Sep 2026 · 0 citations · 76 references
Computer Science

TL;DR

Robion is designed, the first VLA serving and management system for multi-robot, multi-model requests on multi-GPU edge servers that meets SLOs and enables flexible model placements on multi-GPU servers, and integrates an intelligent traffic controller.

Abstract

Vision-Language-Action (VLA) models show high robotic manipulation capabilities via a two-stage design: a Vision-Language Model (VLM) stage followed by an Action Diffusion Transformer (ADiT) stage. Since robots must meet strict Service-Level Objectives (SLOs) for safety, VLA inference is inherently latency-critical. Meeting these SLOs requires high-end GPUs, yet weight, cost, and power constraints preclude integrating such GPUs on-robot. Prior works offload VLA inference to edge servers that serve many robots on VLA models. However, current VLA systems lack support for multi-request, multi-model execution on a multi-GPU server under SLOs, while existing serving systems for multi-stage models are optimized for throughput and stage disaggregation across separate GPUs, which are ill-suited for the millisecond-scale stages of VLA models. We design Robion, the first VLA serving and management system for multi-robot, multi-model requests on multi-GPU edge servers that meets SLOs. Our serving engine disaggregates the VLM and ADiT stages within a GPU via two streams, dynamically restricting the SMs on VLM stream so ADiT always finds SMs to run alongside it, and co-locates multiple models by sharing these streams across them, prioritizing requests by least remaining SLO time. Our management engine enables flexible model placements on multi-GPU servers, and integrates an intelligent traffic controller that maximizes per-model batching under the chosen placement while bounding each GPU's load to meet SLOs. For individual models, Robion serves on average 6.7$\times$ and 1.5$\times$ higher robot load within 98% SLO attainment over vLLM-Omni, the most widely used multi-stage serving system, and Monolithic, which runs VLM and ADiT as a single pipeline, respectively. In a large-scale experiment of serving 8 different models on a 4-GPU server, Robion can serve up to 64 robots within 98% SLO attainment.

View source

Similar papers

Preprint Sep 2026

Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies

Vision-Language-Action (VLA) policies commonly run Vision-Language Model (VLM) backbones with billions of parameters at every policy inference, which costs latency and energy. We revisit a decoupled alternative for multi-task manipulation: separate vision and language encoders whose representations condition a compact...

Xia-Tao Sun, Chen Liang, Zi-Yao Zeng et al. · 2 citations
#machine learning Preprint Sep 2026

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

This work proposes a framework that exposes backbone depth V, action expert depth A, and denoising steps $D$ as three jointly configurable compute axes in a VLA, and introduces a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deep...

Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors whil...

Han-Chu Zhou, De-Chen Gao, Hang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

A symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers is proposed, providing preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain s...

Yong-Shen Zhao, Han Gao, Bao-Ping Cheng et al. · 0 citations
Preprint Sep 2026

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchro...

Bing-Xuan Li, Si-Qi Song, Yi-Zhuo Wu et al. · 0 citations
Preprint Sep 2026

KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models

Vision-Language-Action (VLA) models have emerged as foundational models for next-generation robotics. High VLA inference throughput is critical for meeting the control-rate requirements of robots. VLA models comprise two phases, a vision-language model (VLM) and an action head, that can be decoupled and executed asynch...

A. Li, Christina Giannoula, N. Vijaykumar · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.