This paper aims to reduce latency and the dynamic memory usage of the π 0.5 model by performing timestep distillation on the diffusion process in the π 0.5 action head to create a 3-step diffusion action head.
An open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training.
Mathilde Kappel, Clémence Grislain, Mohamed Chetouani et al.· 0 citations
Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%....
TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
This work instantiates Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies that meets the real-time control requirements of dynamic real-world manipulation, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics.
Perry Dong, Kuo-Han Hung, D. Sadigh et al.· 0 citations
Robion is designed, the first VLA serving and management system for multi-robot, multi-model requests on multi-GPU edge servers that meets SLOs and enables flexible model placements on multi-GPU servers, and integrates an intelligent traffic controller.
Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri et al.· 0 citations
Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts obje...
Caoliwen Wang, Meng-Di Wang, Heng Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.