Skip to content
Preprint

Towards High-DoF Dexterous Manipulation through VLA Post-Training

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

A unified four-step post-training pipeline comprising a learned temporal hand-action codec, supervised fine-tuning, DAgger, and real-world residual reinforcement learning provides a practical path for adapting VLA foundation models to reliable real-world dexterous manipulation.

Abstract

Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly difficult: their broad behavioural repertoire and high degree of freedom create a large and structured action space. Three obstacles are central: open-source VLAs do not natively provide an action interface for high-DoF hands; gesture mismatch during human-gated DAgger takeover creates command discontinuities and contaminates corrective trajectories; and reinforcement learning in the raw joint space is sample-inefficient. We present a unified four-step post-training pipeline comprising a learned temporal hand-action codec, supervised fine-tuning, DAgger, and real-world residual reinforcement learning. The codec adapts a pretrained VLA to absolute dexterous-hand commands. Buffered rollback, pose alignment, and smooth command blending enable continuous, task-relevant DAgger corrections, while latent residual RL confines exploration to coordinated hand motions captured by the codec. We evaluate the pipeline on five diverse real-world tasks spanning bimanual transfer, in-hand reorientation, and tool use. Within the reported post-training budgets, the resulting policies achieve 100\% success on every evaluated task over 20 trials per task. These results provide a practical path for adapting VLA foundation models to reliable real-world dexterous manipulation.

View source

Similar papers

Preprint Aug 2026

Pre-training Visual Dexterity in Simulation

Simulation Pre-training for Dexterity (SPD) is introduced, a pre-training framework for dexterous manipulation that uses data entirely collected in simulation and outperforms training behavior cloning policies from scratch, showing that simulation teleoperation is a viable pre-training source for real-world dexterous m...

Sarthak Kamat, Adam Rashid, Satvik Sharma et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limite...

Zimu Han, Yi-Ming Zeng, Ji-Yao Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning

We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterou...

Jayjun Lee, Jessica Yin, Asif Rana et al. · 2 citations
Preprint Sep 2026

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchro...

Bing-Xuan Li, Si-Qi Song, Yi-Zhuo Wu et al. · 0 citations
Preprint Sep 2026

ForceRFT: Refining VLA Actions through Force-Guided Residual Reinforcement Learning

Force-conditioned vision-language-action (VLA) policies can respond to contact, but when trained solely on demonstrations, their recovery behavior may be limited by demonstration coverage, and they do not learn from deployment outcomes. Human corrective imitation provides additional recovery examples, but its objective...

Yi-Cheng Wang, Chao-Yang Zhang, Xu-Qi Su et al. · 0 citations
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

This work instantiates Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies that meets the real-time control requirements of dynamic real-world manipulation, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics.

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.