This submission to the MedReason 2026 challenge is described, covering multiple-choice (MCQ) and open-ended (OE) medical visual question answering (VQA) under fully offline, containerized inference, and it is found that MCQ retrieval must compare answer \emph{semantics} rather than answer labels.
Abstract
We describe our submission to the MedReason 2026 challenge, covering multiple-choice (MCQ) and open-ended (OE) medical visual question answering (VQA) under fully offline, containerized inference. Our first finding is that MCQ retrieval must compare answer \emph{semantics} rather than answer labels: labels are independently assigned per question, so copying a retrieved neighbor's label transfers no useful information, whereas scoring each current option's text against correct-answer text from similar training cases raises retrieval-only accuracy from 20.0\% to 57.5\% on a 200-case retrieval-excluded development holdout. Our second finding attributes the submitted system's accuracy: holding the task-specific MCQ Low-Rank Adaptation (LoRA) adapter fixed and varying the number \(k\) of in-prompt retrieved examples changes accuracy by at most one case --- 187/200 (93.5\%) at both \(k=0\) and the adapter's training-time \(k=1\), 188/200 (94.0\%) at the packaged runtime's default \(k=3\) --- and the submitted confidence-gated override adds no net accuracy on top of \(k=3\), selecting the VLM in 198/200 cases. With the final MCQ adapter fixed, retrieval changes accuracy by at most one case, and gating provides no net gain. On 20 OE cases, token-F1 and RaTEScore~\cite{zhao2024ratescore} decrease as \(k\) grows, but paired sign tests on token-F1 differences are nonsignificant (\(p \ge 0.29\)); a single-annotator comparison found 6/20 wrong-anchor errors for the final configuration and 14/20 for an earlier configuration that jointly differed in routing, adapter, and prompting. The system reaches 94.0\% MCQ accuracy on the development holdout and 93.20\% on the organizer's official pre-evaluation, versus 29.43\% for the off-the-shelf reference baseline, while both of the organizer's open-ended scores are lower than that baseline's (ground-truth agreement 1.245 versus 1.588, visual accuracy 1.995 versus 2.696, each out of 4).
Medical question answering (QA) plays a crucial role in clinical decision support, yet robust performance requires models to effectively distinguish relevant evidence from topically similar distractors within retrieved contexts. Existing Vietnamese medical QA benchmarks, however, focus exclusively on zero-shot evaluati...
Nhan Phuoc Thanh Tran, P. Huynh, Trương Quốc Tuấn Trương et al.· International Conference on...· 0 citations
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text g...
The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.
Across six medical multimodal MCQ datasets, this work separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key, showing that medical image-reasoning claims require route-level evidence.
Ben Wang, Yi-Fan Zhang, Jia-Qing Yu et al.· 0 citations
Med-VQA systems help alleviate the diagnostic burden on radiologists by providing decision support to clinicians through the interpretation of clinical queries based on radiological images. Traditional models are classification-based and are limited to a restricted response space. This proves insufficient in complex me...
E. Balık, Mehmet Kaya· Automation, Control, and Inf...· 0 citations
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibratio...