Jul 2026· American Journal of Pathology· 0 citations· 37 references
Medicine
TL;DR
Current pathology VLMs support a growing range of use cases, including image-text retrieval, label-efficient classification, visual question answering, abnormality localization, anomaly detection, report generation, and agentic workflow support, according to a review of current systems.
Abstract
Vision-language models (VLMs) represent an emerging class of multimodal artificial intelligence (AI) systems that integrate visual information with natural-language understanding and generation. In computational pathology, VLMs provide a framework for aligning histologic morphology from whole slide images (WSIs) with pathology reports, and other text-based knowledge sources. This review summarizes the technical foundations, major applications, evaluation strategies, and deployment considerations of pathology VLMs. Current pathology VLMs support a growing range of use cases, including image-text retrieval, label-efficient classification, visual question answering, abnormality localization, anomaly detection, report generation, and agentic workflow support. These capabilities are enabled by image encoders, text encoders or large language models, multimodal alignment strategies, and, in some systems, generative language components. Despite rapid progress, several barriers remain. Evaluation of pathology VLMs is constrained by limited domain-specific benchmarks, insufficient assessment of visual grounding, overreliance on text-based metrics, vulnerability to hallucination, and uncertain robustness under data shift. Clinical translation also requires validation across institutions, scanners, staining protocols, tissue types, and patient populations, together with workflow integration, regulatory oversight, data privacy, cybersecurity, and pathologist accountability. VLMs are therefore best viewed as assistive systems that may augment rather than replace pathologists. Responsible development will require close collaboration among pathologists, computational scientists, health systems, and regulatory stakeholders to ensure that VLMs improves pathology practice in a safe, interpretable, and clinically meaningful manner.
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.
Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al.· 0 citations
PathVU is introduced, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology that provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Zongyi Chen, Yuping Liang, Jie Lin et al.· arXiv.org· 2 citations
This Review examines volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction, and introduces a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation.
Zanting Ye, Shengyuan Liu, Xin Liu et al.· 0 citations
A definitive taxonomy of the medical VLM landscape is provided, tracing the evolution from early Contrastive Alignment and Generative MLLMs to the cutting-edge frontiers of Dense Pixel-Grounding, Sparse Mixture-of-Experts (MoE), and Reasoning-Incentivized (RL) architectures.
Taha Razzaq, Murtaza Taj, Asim Iqbal· Journal of Biomedical Inform...· 0 citations
Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meaningful associations between localized visual findings and clinical terminology, and generate diagnostically relevant descriptions. This study proposes an Attention-Guided Vision-Language Model (AG-VLM) for automated radiology report generation that integrates multi-scale visual feature extraction, spatial attention-guided abnormality localization, cross-modal vision-language alignment, and an attention-aware Transformer-based report decoder. The proposed framework selectively emphasizes clinically significant image regions while suppressing redundant background information, thereby strengthening the correspondence between radiographic findings and generated textual descriptions. Experiments were conducted using chest radiograph–report pairs, with performance evaluated using standard natural-language-generation and clinical-consistency measures. The proposed AG-VLM achieved a BLEU-1 score of 0.521, BLEU-2 of 0.387, BLEU-3 of 0.301, BLEU-4 of 0.243, METEOR of 0.286, ROUGE-L of 0.418, and CIDEr of 0.472. For clinical content preservation, the framework obtained a clinical precision of 0.861, recall of 0.842, and F1-score of 0.851. The attention-guided architecture also achieved an abnormality localization accuracy of 91.7% and an overall clinical finding accuracy of 92.4%. Compared with the selected baseline vision-language report-generation model, AG-VLM improved BLEU-4 by 12.5%, METEOR by 9.6%, ROUGE-L by 8.3%, and clinical F1-score by 7.9%. These results indicate that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinically coherent, and contextually relevant radiology reports. The proposed framework therefore provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use.
P. Dayaker, M. Vignesh, I. Z. et al.· International journal of com...· 0 citations
Large language models (LLMs) and vision-language models represent a fundamentally different category of artificial intelligence (AI) compared to prior image analysis approaches in digital pathology, which have largely been based on convolutional neural network architectures. This review from the American Society of Cytopathology Clinical Practice Committee examines the current evidence for LLM and vision-language model applications in cytopathology, including structured reporting, diagnostic assistance, quality control, education, and workflow integration. The distinction between applications with preliminary evidence and those that remain hypothetical is described. A detailed assessment of the challenges that must be addressed before clinical deployment, including hallucination risk, limited explainability, bias, data privacy, validation gaps, and infrastructure barriers is discussed. A review of the regulatory landscape in the United States and European Union as it applies to AI-enabled software as a medical device is provided. Recommendations addressing cytopathology-specific benchmarks, multi-institutional validation, transparent governance, and incremental deployment beginning with low-risk applications are suggested. In the current environment, LLMs have the potential to augment cytopathology practice, but responsible adoption requires rigorous validation and sustained collaboration among cytopathologists, AI researchers, and regulatory bodies.
K. Bilal, Joanna A Gibson, David Kim et al.· Cancer Cytopathology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.