Skip to content
Preprint

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

Sep 2026 · 0 citations · 13 references
Computer Science

TL;DR

An open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training.

Abstract

Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as sequences of action chunks. Training and evaluating these models requires large-scale collections of real-world demonstrations, pairing robot actions with the corresponding visual and proprioceptive observations. Collecting such data on real hardware typically relies on human teleoperation, making the process costly, time-consuming, and difficult to scale. We present an open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training. The same deployment stack is then reused, in closed-loop, to evaluate a trained policy on that setup, so that data collection and evaluation share an identical hardware configuration. Because each real recording is paired with the simulated trajectory that produced it, the protocol also yields a direct measurement of the sim-to-real gap. We release the collected datasets on Hugging Face together with the pipeline source code https://gitlab.isir.upmc.fr/kappel/sim2real_public_chunk_control.

View source

Similar papers

Preprint Aug 2026

DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation

DREAM is presented, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration, and whether it can serve as a scalable data-collection system for the deployment workspace.

Makoto Sato, T. Matsushima, Yutaka Matsuo et al. · 1 citation
Preprint Oct 2026

ChunkVLA-AM: Parallel Action Chunking for Vision-Language-Action Robot Control in Additive Manufacturing

Vision-language-action (VLA) models unify visual perception, language understanding, and action generation, offering new opportunities for automation in additive manufacturing (AM). However, deployment in AM remains challenging because adapting these models to unseen robot embodiments is costly, and performance can deg...

Zhuo Liu, Kai-Chuang Zhang, Jin-Man Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limite...

Zimu Han, Yi-Ming Zeng, Ji-Yao Zhang et al. · 0 citations
Preprint Sep 2026

R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels.

Yi-Di Wang, Fei-Xiang Ruan, Ruo-Qu Chen et al. · 2 citations
Preprint Aug 2026

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI is proposed, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters and seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference.

Zhe Liu, Jinghua Hou, Yuxiang Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.