Parameter-Efficient Fine-Tuning of InstructBLIP for Medical Visual Question Answering
Abstract
Med-VQA systems help alleviate the diagnostic burden on radiologists by providing decision support to clinicians through the interpretation of clinical queries based on radiological images. Traditional models are classification-based and are limited to a restricted response space. This proves insufficient in complex medical scenarios requiring free-text generation. In this study, a generative Med-VQA framework is presented by adapting InstructBLIP, a large-scale visual-language model (VLM), to the medical domain while ensuring low computational cost. The model's adaptation was achieved using the Low-Rank Adaptation (LoRA) technique, which combines 8-bit quantization and Parameter-Efficient Fine-Tuning (PEFT). With this approach, approximately 96.5% of the original parameters were frozen, leaving only ~3.47% (~141.6 million) of the total parameters open for training on standard hardware. The proposed method was evaluated on VQA-RAD, a reference dataset in the field of radiology, using a strict image-based train/test split that prevents data leakage, in the context of both closed-ended and open-ended questions. Experimental results demonstrate that the model successfully aligns with the medical semantic space, achieving 70.51% accuracy and a weighted F1 score of 0.7059 on closed-ended questions expecting a yes/no response. In text generation evaluations for open-ended questions, the BLEU-1 score was measured at 0.2735 and the ROUGE-L score at 0.2580. Qualitative and quantitative findings demonstrate that the LoRA-based adaptation provides a strong foundation for transitioning from classification to generative architectures in the Med-VQA domain; they also indicate that current vulnerabilities in open-ended questions requiring anatomical localization could be addressed in the future through knowledge-enhanced (RAG) architectures.