A lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment is proposed, highlighting its efficiency and potential for clinical deployment.
Abstract
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment. Specifically, our approach features two key components: Frequency-Memory Fusion (FMF), which enhances low-frequency features by retrieving from a learnable memory bank built on DCT decomposition, and Graph-Aware Cross-Attention (GACA), which aligns visual-textual features via cross-attention and refines them through graph-convolutional aggregation. To address data scarcity, we construct SynMed-VQA, a large-scale synthetic dataset comprising over 2 million question-answer pairs across 9 imaging modalities and 10 major organs, generated with GPT-4o. Extensive experiments on SynMed-VQA and three other standard biomedical VQA benchmarks demonstrate that MedFG-VQA achieves competitive or superior performance compared to much larger models while maintaining significantly lower computational costs, highlighting its efficiency and potential for clinical deployment.
A dual-branch frequency-domain fusion module that conditions spectral filtering on the input question, enabling adaptive selection of global low-frequency structure and fine-grained high-frequency detail before reconstructing the spatial representation for answer generation is introduced.
The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.
Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must answer free-form clinical questions about an image. While recent vision-language models (VLMs) achieve promising answer accuracy on this task, clinical adoption also requires the model's internal representations to reflect the visual evidence behind its answers. We propose a simple multi-task fine-tuning recipe that constructs auxiliary grounding and description tasks from an existing VQA dataset with minimal additional annotation: expert-annotated polyp masks are reused directly, while a GI-domain pretrained classifier with Grad-CAM localization provides weak supervision for finding categories that lack ground-truth masks. Three small VLM backbones are fine-tuned with low-rank adaptation under matched VQA-only and multi-task recipes on Kvasir-VQA-x1, and we show consistent accuracy gains together with improved implicit alignment between answer tokens and the relevant image region, evaluated on both in-distribution and out-of-distribution data.
Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh et al.· 0 citations
A multilingual medical VQA benchmark over eight languages is constructed, organized into four representative scenarios that isolate the core capabilities medical VQA requires, and a training-free scenario-aware representation engineering method is proposed, leveraging LVLMs's superior English medical VQA capability to steer non-English representations toward their English counterparts at inference time.
Jingbo Wang, Sendong Zhao, Haochun Wang et al.· 0 citations
A novel multimodal RAG framework tailored for MedVQA is proposed, which leverages multimodal data, including medical images, reports, and generated captions, to provide more accurate clinical answers, and introduces a training paradigm that uses captions as auxiliary supervision, enhancing cross-modal alignment via contrastive learning.
Mai A. Shaaban, M. Zarei, Adnan Khan et al.· 0 citations
Medical image de-identification is critical for artificial intelligence research in the healthcare domain. It requires the detection and localization of Protected Health Information (PHI) burned into medical images. While vision-language models have shown strong visual understanding capabilities, recent studies reveal they suffer from attention dispersion, which correctly localizes but fails to perceive small visual details. Moreover, naive feature selection approaches that discard spatial context destroy positional information essential for accurate bounding box prediction. We propose Salient-Q, a saliency-guided vision-language framework that addresses these challenges through three architectural innovations: (1) a Saliency Module that learns per-token PHI probability and amplifies relevant features through adaptive soft gating; (2) a position-aware scout detector module with explicit 2D positional embeddings that preserves spatial relationships, and (3) a saliency-weighted cross-attention with coordinate supervision that aligns attention centers with ground-truth bounding box centers. We also developed a digital-decayed PHI synthesis pipeline and constructed the Decayed-PHI-50K dataset that was used to fine-tune and evaluate the developed models. Salient-Q outperforms other baseline models in terms of PHI identification and localization. Our code is available at: https://github.com/zjsuper/phi\_deidentification\_vlm.
Sicheng Zhou, Zaifu Zhan, Lei Wu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.