Preprint
Aug 2026
Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
Q-Planning is proposed, which equips a large visuomotor BC policy with a small off-policy Q-function and exploits this asymmetry to enable value-guided action selection at inference and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched.
Varun Giridhar, Anant Khandelwal, Jeremy A. Collins et al.
· 0 citations