Point and Pick: Bounding-Box Conditioned Diffusion Policies and Offline RL for Target-Specific Robot Manipulation
The main hypothesis was that a lightweight task-specific diffusion policy, explicitly grounded by a visual bounding-box prompt, could outperform text-only VLA prompting for fine-grained target selection, and that offline RL could further improve robustness beyond imitation learning (IL).
R. Garreta, Joshua Bowden
· 0 citations