Skip to content
Review Open access

RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data

Sep 2026 · Machine Learning and Knowledge Extraction · 0 citations · 13 references

Abstract

Public surgical segmentation datasets contain information about the operative view that extends beyond their pixel labels. RoboSurg-VQA organises this information as a closed-set visual question answering (VQA) benchmark while retaining the source of each answer. It contains 5632 frames from 29 EndoVis 2017/2018 source sequences and 11 task units, with answers derived from source metadata, mask measurements, or machine-generated candidate labels. Seven visual attributes were examined in a blinded 250-frame audit by two study-team reviewers working independently, with disagreements resolved by a third reviewer. We trained a shared model with frozen BiomedCLIP encoders and a 13-answer vocabulary to examine the contributions of image and question inputs. Against mask-derived and candidate labels on held-out sequences, the model achieved a seven-task mean fixed-label Macro-F1 of 0.540. Answer frequency and task-restricted image-only prediction scored 0.391 and 0.400, respectively. With unseen paraphrases, the score was 0.528. On 40 audited frames from held-out sequences, mean Macro-F1 across bleeding, image quality, and glare was 0.535 against the human reference, compared with 0.450 for answer frequency, a gain of 0.085 driven by bleeding. The audit identified attribute-specific label errors, with the greatest reviewer disagreement for glare.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.