Skip to content

ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

Sep 2026 · 0 citations
Computer Science

TL;DR

ReWeight is introduced, a framework that incorporates human data into VLA post-training through demonstration-level retrieval and sample-level weighting and learns a cross-embodiment visuomotor representation that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations.

Abstract

Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data is costly. Egocentric human demonstrations provide a scalable alternative, but directly mixing human and robot data can introduce cross-embodiment discrepancies and degrade policy performance. To address this challenge, we introduce ReWeight, a framework that incorporates human data into VLA post-training through demonstration-level retrieval and sample-level weighting. ReWeight learns a cross-embodiment visuomotor representation that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations. Based on optimal transport, it retrieves human demonstrations relevant to the target robot data and assigns larger weights to samples with smaller cross-embodiment discrepancies. We evaluate ReWeight using $\pi_{0.5}$ across eight simulation tasks and four real-world tasks under both clean and randomized settings. In simulation, ReWeight improves the average success rate of post-trained $\pi_{0.5}$ from 39% with only robot data and 44% with randomly mixed human-robot data to 57%. In the physical experimental setting, it achieves an average success rate of 68.8%, outperforming the baselines by 28.8% and 13.8%, respectively. Overall, ReWeight provides an effective paradigm for transforming abundant egocentric human experience into transferable supervision for robot learning. (Project webpage: https://reweight-vla.github.io/)

View source

Similar papers

Preprint Sep 2026

How to Better Train VLAs: Lessons Learned From the REAL-I Challenge at ICRA 2026

How can robot policies learn more effectively from a fixed demonstration budget? The first Real-world Embodied AI Learning (REAL-I) Challenge at ICRA 2026 examined this question through simulation, real-robot evaluation, and an on-site final on a shared dual-arm humanoid platform. We describe the challenge tasks, data...

Jia-Ming Wang, Ji-Zhuo Chen, Di-Wen Liu et al. · 0 citations
Preprint Sep 2026

Rho: A Foundation for Efficiently Adaptable VLA Models

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments re...

Rho Team Simran Bagaria, Daphne Chen, Dean Fortier et al. · 0 citations
Preprint Sep 2026

RoboDrop: Curating VLA Post-Training Data via Local Gradient Compatibility

Vision--language--action (VLA) models acquire broad generalization through large-scale pretraining, yet adapting them to a new task and robot embodiment still requires post-training on newly collected data. Unlike pretraining, post-training targets task- and embodiment-specific adaptation, making it particularly sensit...

Run-Ze Xu, Yuan-Fan Xu, Cui-Jie Xu et al. · 0 citations
Preprint Sep 2026

Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment Transfer

The performance of learned robot visuomotor policies depends heavily on the size and quality of their training data, yet collecting high-quality demonstrations remains costly for robots in the real world. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action spa...

Yi-Qin Wang, Mrinal Verghese, Jeff G. Schneider · 0 citations
Preprint Aug 2026

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-d...

Sen-Qiao Yang, Cheng-Yao Wang, Yuxin Chen et al. · 4 citations
#machine learning Preprint Sep 2026

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately. In this work, we systemati...

Jinho Jeong, Se June Joo, Jaehyun Kang et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.