PROBE: Manipulation-Grounded Visual Question Answering with VLM Agents
This work develops PROBE, a framework for benchmarking and finetuning VLM agents on cluttered tabletop scenes, and designs PROBE-Agent, a finetuning recipe to distill successful trajectories from a powerful teacher foundation model to a smaller open-weight model using a mixed data recipe that encourages manipulation-efficient question answering.