Skip to content
Preprint

Dense to MoE Adaptation for Compact Vision Language Action Policies

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

The results suggest that dense to MoE adaptation with dynamic expert deactivation is a practical direction for reducing active VLA model size without severe performance loss.

Abstract

Vision language action (VLA) policies continue to grow in parameter count, making deployment on resource-constrained robot platforms difficult. The central goal is to reduce the number of LLM-side parameters retained in the deployed policy while preserving downstream task performance. Our approach, AdaDE, adapts selected dense feed forward blocks into mixture of experts (MoE) layers and derives expert retention masks from router statistics during fine tuning. The Dense2MoE conversion preserves the original dense FFN function at initialization, so expert deactivation can start without a separate recovery stage. Instead of using a fixed shutdown rule, expert masks are updated dynamically from router usage statistics, with staged training and expert protection to avoid early collapse. With 40% of the LLM parameters deactivated, AdaDE retains 95.7% average success in LIBERO and 42.0% average success across all 50 RobotWin2.0 tasks. These results suggest that dense to MoE adaptation with dynamic expert deactivation is a practical direction for reducing active VLA model size without severe performance loss.

View source

Similar papers

Preprint Sep 2026

Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies

Vision-Language-Action (VLA) policies commonly run Vision-Language Model (VLM) backbones with billions of parameters at every policy inference, which costs latency and energy. We revisit a decoupled alternative for multi-task manipulation: separate vision and language encoders whose representations condition a compact...

Xia-Tao Sun, Chen Liang, Zi-Yao Zeng et al. · 2 citations
#small language model Preprint Sep 2026

MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?

Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed...

Kohei Sendai, T. Matsushima, Yusuke Iwasawa · 3 citations · ⚡1
#machine learning Preprint Sep 2026

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their ver...

Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al. · 0 citations
Preprint Sep 2026

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%....

Chiyoung Kim, S. Choi, Minhyeok Lee · 0 citations
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, cre...

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget...

Zhi-Peng Tang, Xin-Da Chen, Wei-Ning Rao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.