Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while groundin...