Skip to content

HiVLR: Hierarchical Vision-Language Reasoning for interpretable zero-shot radiography image understanding

Jul 2026 · Medical Image Anal. · Vol 113, pp. 104189 · 0 citations · 24 references
Medicine Computer Science

TL;DR

This work revisits how human doctors reason from a patient's radiology image for diagnosis, and proposes a Hierarchical Vision-Language Reasoning (HiVLR) framework based on the clinical diagnostic workflow, and attaches a concept-based interpretable diagnosis block to improve the accuracy and interpretability in downstream tasks simultaneously.

Abstract

Medical vision-language pre-training on image-report pairs has shown great potential to facilitate downstream image understanding tasks. However, prior approaches commonly exhibited limited accuracy on zero-shot image tasks, and lacked sufficient interpretability for unseen disease diagnosis, posing substantial usability concerns and trust issues for safety-critical medical applications. To alleviate them, we revisit how human doctors reason from a patient's radiology image for diagnosis, and propose a Hierarchical Vision-Language Reasoning (HiVLR) framework based on the clinical diagnostic workflow. In specific, we structure feature investigation into sequential rounds of thinking, i.e., (1) spotting the suspicious pathology observations (e.g., obscure) from all visual and textual inputs first and then (2) determining possible diagnostic findings (e.g., pneumonia) that match all pathology observations, to derive accurate disease predictions without compromising transparency in model decision making. Each round of thinking needs to analyze the inputted visual and textual embeddings by coarsely aligning them with cross-attention to establish global correspondences, highlight region-level visual features containing specific clinical content by prompt tuning-enabled fine-grained filtering, and then interpret the visual features in a condensed understanding to derive diagnostic-pertinent discoveries. Importantly, we enforce concept-level cross-modal compliance by ensuring that visual and textual features corresponding to the same clinical content are semantically consistent across concept dimensions (e.g., texture, shape, border). Based on this, we attach a concept-based interpretable diagnosis block to improve the accuracy and interpretability in downstream tasks simultaneously. Experiments showed that our approach greatly outperformed competing approaches on diverse zero-shot image tasks with superior interpretability.

View source

Similar papers

Jul 2026

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

RadSight is proposed, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures that achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable c...

Jianqi Liu, Weiwei Cao, Wan-Xing Chang et al. · 0 citations
Open access Aug 2026

Medrecord-CLIP: enhancing fundus disease diagnosis via EHR-guided vision-language pre-training

This work proposes MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations that highlights the critical value of integrating personalized clinical context to enhance the generaliza...

Lei Shi, Wenbin Zhai, Lei Yu et al. · 0 citations
Jul 2026

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

PathVU is introduced, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology that provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.

Zongyi Chen, Yu-Ping Liang, Jie Lin et al. · 2 citations
Open access Aug 2026

Attention-Guided Vision-Language Model for Automated Radiology Report Generation

The proposed AG-VLM framework provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use and indicates that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinical...

P. Dayaker, M. Vignesh, I. Z. et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

A clinically curated Pan-Asia WSI--report dataset is introduced and the REG 2025 benchmark is established as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology model...

Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A visual large language foundational model for medical image recognition using clinician-oriented social media

A FOundational LLM Trained on ThoughtMed-1M (FOLTMed), a scalable paradigm for advancing research on clinically grounded multimodal LLMs, achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M...

Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.