FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
FineART is presented, a densely annotated bimanual manipulation dataset comprising 40,543 episodes and 533,913 subtasks across 151 tasks and FineART-VLA is introduced, a vision-language-action policy that predicts its own next subtask to guide its actions.
Jade Choghari, Pepijn Kooijmans, Mansi Agarwal et al.
· 0 citations