RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data
Abstract
Public surgical segmentation datasets contain information about the operative view that extends beyond their pixel labels. RoboSurg-VQA organises this information as a closed-set visual question answering (VQA) benchmark while retaining the source of each answer. It contains 5632 frames from 29 EndoVis 2017/2018 source sequences and 11 task units, with answers derived from source metadata, mask measurements, or machine-generated candidate labels. Seven visual attributes were examined in a blinded 250-frame audit by two study-team reviewers working independently, with disagreements resolved by a third reviewer. We trained a shared model with frozen BiomedCLIP encoders and a 13-answer vocabulary to examine the contributions of image and question inputs. Against mask-derived and candidate labels on held-out sequences, the model achieved a seven-task mean fixed-label Macro-F1 of 0.540. Answer frequency and task-restricted image-only prediction scored 0.391 and 0.400, respectively. With unseen paraphrases, the score was 0.528. On 40 audited frames from held-out sequences, mean Macro-F1 across bleeding, image quality, and glare was 0.535 against the human reference, compared with 0.450 for answer frequency, a gain of 0.085 driven by bleeding. The audit identified attribute-specific label errors, with the greatest reviewer disagreement for glare.