Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 18 references
Abstract
In radiology, multimodal vision-language models (VLMs) are increasingly used for clinical tasks such as report generation and visual question answering. Their adoption has raised important concerns regarding the transparency and trust-worthiness of the generated text reports, as the reasoning process leading to a given report remains opaque for clinicians. In this work, we investigate how post-hoc attribution methods behave in large generative volumetric VLMs that jointly process 3D scans and clinical text, a setting that remains largely unexplored. We propose to use Post-hoc gradient-based attribution that directly links small changes in the input volume to changes in the model's output probability. The faithfulness and spatial specificity of four attribution methods are subsequently characterized and compared when applied to Med3DVLM, a recent medical VLM for image-text understanding. To enable differentiable attribution at inference time, we apply teacher forcing exclusively during the attribution forward pass so that the stochastic generation path is converted into a differentiable computation graph. The cumulative log-likelihood of the generated response is used as the scalar attribution target. Input-level methods, including Saliency and Integrated Gradients, alongside feature-level techniques such as Grad-CAM and Guided Grad-CAM, are computed on volumetric inputs across diverse clinical question types. Qualitative assessment and a quantitative voxel deletion protocol indicate that input-space gradient methods, especially Integrated Gradients, produce spatially selective and causally faithful relevance maps, whereas feature-level methods like Layer Grad-CAM exhibit significant spatial diffusion and fail to provide discriminative localization in multimodal architectures. These results demonstrate that input-level attributions exhibit a significantly steeper drop in model confidence compared to feature-level methods, confirming they are more faithful and spatially precise, while intermediate-layer maps remain diffuse and non-discriminative.
Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic fluency without addressing the underlying opacity of the visual model. With the emergence of Kolmogorov-Arnold Networks (KANs), whose spline-based components provide inherently interpretable functional units, we investigate whether this architectural transparency can be leveraged to produce more trustworthy textual explanations. We introduce KANEx, the first ever framework that leverages the symbolic transparency of KANs to ground VLM reasoning. This interpretability also made it possible to design KAN-Map, a novel heatmap generation method derived directly from KAN models rather than gradient approximations. We feed these grounded contexts into downstream VLMs for enhanced explainability. Benchmarked on the MIMIC-CXR dataset, we demonstrate that KAN-based architectures with ResNet/ViT baselines demonstrate improved semantic similarity while producing significantly more faithful saliency maps. KAN architectures improve visual localization and downstream reasoning quality by 10%. Our findings suggest that grounding linguistic explanations and visual attributions in mathematically interpretable units is a necessary step toward trustworthy medical AI.
Krithi Shailya, Ananya Lakshmi Ravi, V. VenkatanathanK. et al.· arXiv.org· 0 citations
A large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm that unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training is introduced.
Biomedical vision–language models increasingly support image-grounded clinical dialogue, yet most deployable systems still depend on autoregressive language generation. Such systems tend to truncate answers, react poorly to length instructions, and offer no principled way to signal uncertainty when image evidence is weak. We present MedDiffVL, a biomedical vision-language model that pairs a masked language diffusion backbone with a SigLIP-2 visual encoder and a multimodal alignment pipeline that injects modality and question-type cues. Three inference-time mechanisms target the failure modes of diffusion-based generators in the clinical setting. An adaptive confidence-guided remasking rule uses a time-aware threshold and a short-window stability check to remove repetitive low-quality candidates. A clinically aware length controller selects a target length from question type, modality, and an internal uncertainty estimate. A reliability gate combines visual-evidence and answer-confidence scores to emit, hedge, or escalate a response. On VQA-RAD, SLAKE, and PathVQA, the model reaches 85.42, 92.78, and 94.91% closed-form accuracy and an overall conversation score of 53.42 against a fixed reference. Token repetition falls from 0.18 to 0.06. An ECE falls from 0.137 to 0.034, but this reflects an ECE-surrogate training loss and is not independently validated. These gains are not uniform. The closed-form gains over the prior diffusion model lie within run-to-run variance, and latency stays higher than autoregressive baselines. The main contribution is controllability and reliability-aware decoding, not higher closed-form accuracy. The results indicate that confidence-guided masked diffusion with reliability-aware decoding is a useful direction for controllable and reliability-aware clinical assistants.
Saqib Qamar, Goram Mufarah M. Alshmrani· Technologies· 0 citations
This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-of-Thought.
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.
Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al.· 0 citations
Medical report generation aims to reduce the reporting burden of radiologists by translating medical images into clinically meaningful descriptions. Although generative adversarial networks (GANs) have shown potential for image captioning and report generation, their application to radiology reports remains challenging because radiology text requires standardized negative expressions, adversarial training may be unstable, and lexical metrics alone cannot fully reflect clinical utility. To address these issues, we propose ConCapGAN, a GAN-based medical report generation framework with a visual manifold graph (VMG), a zero-centered Wasserstein regularizer, and a language-style constraint. The VMG explicitly models object–predicate associations between detected image regions and report semantics, thereby helping represent clinically important positive and negative findings. The zero-centered regularizer improves the local stability of adversarial optimization under clearly stated assumptions, and the language-style constraint encourages reports to follow clinician-like phrasing without introducing an additional large language module. Experiments on IU X-Ray, MIMIC-CXR, and LGK show that ConCapGAN achieves competitive natural language generation performance and improved clinical efficacy (as measured by several diagnostic metrics) compared with recent report generation methods.
Yuan Wang, Shijie Xu, Kun Zhou et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.