Skip to content

Boltzmann-driven dynamic annealing network for knowledge-guided radiology report generation

Jul 2026 · Multimedia Systems · Vol 32 · 0 citations · 57 references
Computer Science

TL;DR

A novel Boltzmann-driven Dynamic Annealing Network for Knowledge-guided Radiology Report Generation that employs Boltzmann distribution with an annealing mechanism to process global visual features, while integrating medical knowledge with dynamic sparse attention for precise lesion identification is proposed.

View source

Similar papers

Open access Aug 2026

Attention-Guided Vision-Language Model for Automated Radiology Report Generation

Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meaningful associations between localized visual findings and clinical terminology, and generate diagnostically relevant descriptions. This study proposes an Attention-Guided Vision-Language Model (AG-VLM) for automated radiology report generation that integrates multi-scale visual feature extraction, spatial attention-guided abnormality localization, cross-modal vision-language alignment, and an attention-aware Transformer-based report decoder. The proposed framework selectively emphasizes clinically significant image regions while suppressing redundant background information, thereby strengthening the correspondence between radiographic findings and generated textual descriptions. Experiments were conducted using chest radiograph–report pairs, with performance evaluated using standard natural-language-generation and clinical-consistency measures. The proposed AG-VLM achieved a BLEU-1 score of 0.521, BLEU-2 of 0.387, BLEU-3 of 0.301, BLEU-4 of 0.243, METEOR of 0.286, ROUGE-L of 0.418, and CIDEr of 0.472. For clinical content preservation, the framework obtained a clinical precision of 0.861, recall of 0.842, and F1-score of 0.851. The attention-guided architecture also achieved an abnormality localization accuracy of 91.7% and an overall clinical finding accuracy of 92.4%. Compared with the selected baseline vision-language report-generation model, AG-VLM improved BLEU-4 by 12.5%, METEOR by 9.6%, ROUGE-L by 8.3%, and clinical F1-score by 7.9%. These results indicate that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinically coherent, and contextually relevant radiology reports. The proposed framework therefore provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use.

P. Dayaker, M. Vignesh, I. Z. et al. · 0 citations
Open access Aug 2026

KDMG: knowledge-enhanced dynamic memory and gated fusion for chest X-ray report generation

In order to overcome the challenges of incorrect medical terminology application and inaccurate descriptions in the generated reports due to the static character of medical knowledge and rough features combination, this paper will introduce a novel method in the form of the KDMG model of chest X-ray report generation to reduce these difficulties. KDMG has two main innovations: 1) the Knowledge-Enhanced Dynamic Memory module, which constructs a structured medical database and uses confidence-aware querying and momentum update mechanisms to facilitate dynamic, personalized retrieval of medical knowledge based on static priors; 2) the Gated Residual Feature Fusion module, which emulates the reasoning process of doctors who first view images and subsequently make a judgment, incorporating knowledge into visual features using a spatially adaptable gated network. Experiments on the IU X-Ray and MIMIC-CXR datasets demonstrate that KDMG achieves near state-of-the-art performance across both text generation metrics, such as BLEU-4 and CIDEr, and clinical accuracy metrics, including CheXpert F1 Score. Furtherrmore, ablation studies and fusion strategy comparisons further validate the effectiveness and design rationale of each module.

Jie Xiong, Chang-Fa Wei, Hui-Na Liu · 0 citations
Jul 2026

Pathologist Attention-Aligned Report Generation for Prostate Histopathology

The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracted from whole-slide images (WSIs). Human attention helps medical imaging tasks such as classification and segmentation, and becomes a strong semantic cue for identifying diagnostically informative regions for report generation. In this paper, we introduce human attention into the training of pathologist report generation models. To this end, we collected a multimodal human-attention dataset of 121 prostate WSIs annotated with pathologists'multi-scale viewport trajectories synchronized with the pathologists'verbal descriptions and cursor movements for five clinically relevant components (e.g., Gleason patterns). Using this dataset, we finetune two report generation models with an attention-alignment loss that regularizes the model attention over image patches to match the distribution of pathologist attention. We evaluate our approach on prostate cancer report generation and visual question answering using two models with different internal attention mechanisms (i.e., how image tokens are integrated into the language decoder). Experiments show average gains of 10.9% on NLP-based metrics and 19.3% in accuracy across five clinically relevant report components. Further, model attention maps extracted at inference time, with minimal computational overhead, align more closely with pathologist attention, providing stronger visual support for the generated reports by highlighting the regions that most influence the output.

Ruo-Yu Xue, S. Singh, Souradeep Chakraborty et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.

Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.