Jul 2026· Medical Image Anal.· Vol 113, pp.
104189
· 0 citations· 24 references
MedicineComputer Science
TL;DR
This work revisits how human doctors reason from a patient's radiology image for diagnosis, and proposes a Hierarchical Vision-Language Reasoning (HiVLR) framework based on the clinical diagnostic workflow, and attaches a concept-based interpretable diagnosis block to improve the accuracy and interpretability in downstream tasks simultaneously.
Abstract
Medical vision-language pre-training on image-report pairs has shown great potential to facilitate downstream image understanding tasks. However, prior approaches commonly exhibited limited accuracy on zero-shot image tasks, and lacked sufficient interpretability for unseen disease diagnosis, posing substantial usability concerns and trust issues for safety-critical medical applications. To alleviate them, we revisit how human doctors reason from a patient's radiology image for diagnosis, and propose a Hierarchical Vision-Language Reasoning (HiVLR) framework based on the clinical diagnostic workflow. In specific, we structure feature investigation into sequential rounds of thinking, i.e., (1) spotting the suspicious pathology observations (e.g., obscure) from all visual and textual inputs first and then (2) determining possible diagnostic findings (e.g., pneumonia) that match all pathology observations, to derive accurate disease predictions without compromising transparency in model decision making. Each round of thinking needs to analyze the inputted visual and textual embeddings by coarsely aligning them with cross-attention to establish global correspondences, highlight region-level visual features containing specific clinical content by prompt tuning-enabled fine-grained filtering, and then interpret the visual features in a condensed understanding to derive diagnostic-pertinent discoveries. Importantly, we enforce concept-level cross-modal compliance by ensuring that visual and textual features corresponding to the same clinical content are semantically consistent across concept dimensions (e.g., texture, shape, border). Based on this, we attach a concept-based interpretable diagnosis block to improve the accuracy and interpretability in downstream tasks simultaneously. Experiments showed that our approach greatly outperformed competing approaches on diverse zero-shot image tasks with superior interpretability.
RadSight is proposed, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures that achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable c...
This work proposes MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations that highlights the critical value of integrating personalized clinical context to enhance the generaliza...
Lei Shi, Wenbin Zhai, Lei Yu et al.· Health Information Science a...· 0 citations
PathVU is introduced, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology that provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Zongyi Chen, Yu-Ping Liang, Jie Lin et al.· arXiv.org· 2 citations
The proposed AG-VLM framework provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use and indicates that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinical...
P. Dayaker, M. Vignesh, I. Z. et al.· International journal of com...· 0 citations
A clinically curated Pan-Asia WSI--report dataset is introduced and the REG 2025 benchmark is established as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology model...
Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al.· 0 citations
A FOundational LLM Trained on ThoughtMed-1M (FOLTMed), a scalable paradigm for advancing research on clinically grounded multimodal LLMs, achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M...
Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.