RadSight is proposed, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures that achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding.
Abstract
Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent errors in visual interpretation. To systematically trace these failures, we traverse the hierarchy from high-level clinical tasks down to fundamental visual perception. We therefore introduce Perception-Bench, a large-scale benchmark comprising 1.13 million samples that assesses medical MLLMs across six dimensions: attribute judgment, spatial grounding, spatial understanding, disease prediction, anomaly detection, and report generation, spanning both 2D and 3D radiology images. Our analysis on Perception-Bench reveals that existing MLLMs lack the ability to capture even the most basic lesion attributes, such as location, size, and density. This inability to ground clinical outputs in primary visual evidence reveals that the models'diagnostic unreliability is rooted in a critical but overlooked bottleneck in low-level visual perception. Motivated by this, we propose RadSight, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures. RadSight formulates medical image understanding as a four-stage progressive process: visual-language alignment, fine-grained visual perception, clinical diagnosis, and diagnostic interpretation. The model is trained on an 8.37 million perception-oriented corpus using progressive curriculum learning. On Perception-Bench, RadSight consistently outperforms existing MLLMs across all six evaluation dimensions, with particularly strong gains in spatial grounding and clinical diagnosis. It also achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding. Code and model will be publicly available.
The exponential growth in medical imaging volumes necessitates scalable, reliable diagnostic support systems capable of augmenting clinical workflows. This article presents a systematic quantitative evaluation of state-of-the-art Multimodal Large Language Models (MLLMs) for radiology Visual Question Answering (VQA), a task requiring integrated visual perception and clinical reasoning. We benchmark five leading models — GPT5-Nano, Gemini 3 Flash, Qwen3-VL-8B, LLaVA Next, and Llama 3.2 Vision — on the VQA-RAD dataset under a rigorous zero-shot protocol with standardized prompts and comprehensive precision–recall–F1 evaluation. Our empirical analysis reveals that Gemini 3 Flash achieves superior balanced performance (F1 = 0.78, Accuracy = 0.78, Recall = 0.83), while Qwen3-VL-8B attains the highest precision (0.78) while also maintaining competitive recall. These outcomes demonstrate that general-purpose MLLMs can perform competitively with specialized medical models in tasks such as modality and organ recognition, but still struggle with abnormality detection and complex clinical reasoning. The findings reinforce that MLLMs currently serve best as assistive decisionsupport tools rather than autonomous diagnostic agents, and highlight the potential of retrieval-augmented and context-aware strategies for improving clinical reliability and interpretability.
Cristovão Pessoa Cândido, Matheus Alves de Oliveira Lima, C. de Souza Baptista et al.· International Journal of Sem...· 0 citations
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.
Baochen Fu, Wenzhi Deng, Baihao Jin et al.· arXiv.org· 2 citations
PathVU is introduced, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology that provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Zongyi Chen, Yuping Liang, Jie Lin et al.· arXiv.org· 2 citations
A large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm that unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training is introduced.
This Review examines volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction, and introduces a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation.
Zanting Ye, Shengyuan Liu, Xin Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.