Skip to content

Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA

Sep 2026 · 0 citations · 14 references
Computer Science

TL;DR

This submission to the MedReason 2026 challenge is described, covering multiple-choice (MCQ) and open-ended (OE) medical visual question answering (VQA) under fully offline, containerized inference, and it is found that MCQ retrieval must compare answer \emph{semantics} rather than answer labels.

Abstract

We describe our submission to the MedReason 2026 challenge, covering multiple-choice (MCQ) and open-ended (OE) medical visual question answering (VQA) under fully offline, containerized inference. Our first finding is that MCQ retrieval must compare answer \emph{semantics} rather than answer labels: labels are independently assigned per question, so copying a retrieved neighbor's label transfers no useful information, whereas scoring each current option's text against correct-answer text from similar training cases raises retrieval-only accuracy from 20.0\% to 57.5\% on a 200-case retrieval-excluded development holdout. Our second finding attributes the submitted system's accuracy: holding the task-specific MCQ Low-Rank Adaptation (LoRA) adapter fixed and varying the number \(k\) of in-prompt retrieved examples changes accuracy by at most one case --- 187/200 (93.5\%) at both \(k=0\) and the adapter's training-time \(k=1\), 188/200 (94.0\%) at the packaged runtime's default \(k=3\) --- and the submitted confidence-gated override adds no net accuracy on top of \(k=3\), selecting the VLM in 198/200 cases. With the final MCQ adapter fixed, retrieval changes accuracy by at most one case, and gating provides no net gain. On 20 OE cases, token-F1 and RaTEScore~\cite{zhao2024ratescore} decrease as \(k\) grows, but paired sign tests on token-F1 differences are nonsignificant (\(p \ge 0.29\)); a single-annotator comparison found 6/20 wrong-anchor errors for the final configuration and 14/20 for an earlier configuration that jointly differed in routing, adapter, and prompting. The system reaches 94.0\% MCQ accuracy on the development holdout and 93.20\% on the organizer's official pre-evaluation, versus 29.43\% for the off-the-shelf reference baseline, while both of the organizer's open-ended scores are lower than that baseline's (ground-truth agreement 1.245 versus 1.588, visual accuracy 1.995 versus 2.696, each out of 4).

View source

Similar papers

Conference Aug 2026

Retrieval-Augmented Fine-Tuning with Reasoning Distillation for Vietnamese Medical Question Answering

Medical question answering (QA) plays a crucial role in clinical decision support, yet robust performance requires models to effectively distinguish relevant evidence from topically similar distractors within retrieved contexts. Existing Vietnamese medical QA benchmarks, however, focus exclusively on zero-shot evaluati...

Nhan Phuoc Thanh Tran, P. Huynh, Trương Quốc Tuấn Trương et al. · 0 citations
#natural language process... Preprint Sep 2026

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text g...

E. Nourbakhsh, Ke Yang, Anthony Rios · 0 citations
Open access Aug 2026

Retrieval-augmented generation for medical question answering: a multi-metric performance evaluation

The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.

Yunus Kökver · 0 citations
Preprint Sep 2026

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Across six medical multimodal MCQ datasets, this work separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key, showing that medical image-reasoning claims require route-level evidence.

Ben Wang, Yi-Fan Zhang, Jia-Qing Yu et al. · 0 citations
Conference Sep 2026

Parameter-Efficient Fine-Tuning of InstructBLIP for Medical Visual Question Answering

Med-VQA systems help alleviate the diagnostic burden on radiologists by providing decision support to clinicians through the interpretation of clinical queries based on radiological images. Traditional models are classification-based and are limited to a restricted response space. This proves insufficient in complex me...

E. Balık, Mehmet Kaya · 0 citations
#artificial intelligence Preprint Sep 2026

ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering

In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibratio...

Yu-Yang Chen · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.