Jul 2026· IEEE journal of biomedical and health informatics· Vol PP, pp. 1-10· 0 citations· 45 references
Computer ScienceMedicine
TL;DR
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models, designed to reduce potential biases in data and to enable evaluation of the models' ability to handle out-of-distribution inputs, which were difficult in existing datasets.
Abstract
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models have demonstrated significant potential in the field of medical vision-language tasks. However, it is becoming increasingly clear that evaluating these models' performance, whether they are applied to natural or medical images, is challenging. The critical question is whether the models can accurately understand an input image while associating it with relevant input text. To address this, Medical-Checklist imposes a binary test on the models: they are given an image and two captions, where one is correct and the other incorrect, and the model must select the correct one. The incorrect caption contains a single medical concept (word or phrase) that is inaccurately substituted from the correct caption. Although the task is simple, this simplicity enables the unified assessment of diverse multimodal models designed and learned on different principles. It also enables us to verify whether models correctly understand a wide range of medical concepts across various medical sub-domains. Medical-Checklist is designed to reduce potential biases in data and to enable evaluation of the models' ability to handle out-of-distribution inputs, which were difficult in existing datasets. When evaluating four state-of-the-art medical multimodal models with Medical-Checklist, it was revealed that despite their excellent performance in specific tasks such as Med-VQA, they may not correctly understand images, suggesting a long journey ahead for clinical application. The dataset and code will be made public upon acceptance.
RadSight is proposed, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures that achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding.
In order to produce meaningful textual interpretations of intricate clinical images, medical image captioning has become a significant field of study at the nexus of computer vision and natural language processing. Due to the lack of explicit modeling of medical entities, current methods frequently fail to produce descriptions that are both semantically valid and clinically useful, despite notable advances in deep learning and vision language modeling. Typical captioning methods in particular, fall short of being able to retrieve fine grained diagnostic information and maintain semantic consistency with clinical findings as they focus on global features. This paper addresses these limitations by presenting an entity-aware medical image captioning approach which aims to identify and incorporate clinically relevant entities into the caption generation, including but not limited to, diagnostic finding, anatomical structures, or diagnostic characteristics. The proposed method utilizes entity level representations as a means of guiding the captioning process thereby ensuring a tighter semantic consistency between visual modalities and the resultant textual output. Consequently, this leads to more comprehensible, informative and clinically relevant generated reports. Additionally, inclusion of entity awareness can aid the model in effectively understanding the relationships between medical concepts leading to captions more consistent with medical expertise. The results demonstrate that explicit modeling with structured semantic information within vision-language frameworks are crucial and that entity-aware methods have the potential to greatly improve captioning. This work has the ability to help advance health intelligence applications which will serve to better assist clinical decision making, scale the processing of medical images, and facilitate accurate medical documentation.
Shaik Rafi, Syed Rizwana, P. Drutika et al.· IEEE Access· 0 citations
The exponential growth in medical imaging volumes necessitates scalable, reliable diagnostic support systems capable of augmenting clinical workflows. This article presents a systematic quantitative evaluation of state-of-the-art Multimodal Large Language Models (MLLMs) for radiology Visual Question Answering (VQA), a task requiring integrated visual perception and clinical reasoning. We benchmark five leading models — GPT5-Nano, Gemini 3 Flash, Qwen3-VL-8B, LLaVA Next, and Llama 3.2 Vision — on the VQA-RAD dataset under a rigorous zero-shot protocol with standardized prompts and comprehensive precision–recall–F1 evaluation. Our empirical analysis reveals that Gemini 3 Flash achieves superior balanced performance (F1 = 0.78, Accuracy = 0.78, Recall = 0.83), while Qwen3-VL-8B attains the highest precision (0.78) while also maintaining competitive recall. These outcomes demonstrate that general-purpose MLLMs can perform competitively with specialized medical models in tasks such as modality and organ recognition, but still struggle with abnormality detection and complex clinical reasoning. The findings reinforce that MLLMs currently serve best as assistive decisionsupport tools rather than autonomous diagnostic agents, and highlight the potential of retrieval-augmented and context-aware strategies for improving clinical reliability and interpretability.
Cristovão Pessoa Cândido, Matheus Alves de Oliveira Lima, C. de Souza Baptista et al.· International Journal of Sem...· 0 citations
Current evidence supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.