Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 16 references
Abstract
MedQwen-VGR1 is a novel multimodal visionlanguage model (VLM) trained for medical visual question answering (Med-VQA), longitudinal temporal diagnosis, and drug interaction analysis across radiology and pathology imagery. The model undergoes a multi-stage training pipeline comprising: (1) Continuous Domain-Adaptive Pretraining on heterogeneous medical image-text corpora; (2) Supervised FineTuning on expert-annotated multimodal datasets spanning sequential imaging conversations; (3) Human Preference Alignment via Direct Preference Optimization (DPO) and Group Relative Preference Optimization (GRPO) applied to Vision-Guided Chain-of-Thought (CoT) Reasoning trajectories; and (4) Smart Memory Module integration enables Cross-Visit Context Retention and Comparative Reasoning on sequential or temporal pathologies. MedQwen-VGR1 processes sequential medical images and textual queries to generate stepwise, evidence-grounded rationales correlating visual features across timepoints, predict adverse drug interactions from images/metadata, and output calibrated reliability scores. Evaluations on PathVQA, SLAKE, and VQA-RAD demonstrate superior performance over LLaVA-Med (+14.7% accuracy) and BioViL-T (+12.3% temporal reasoning), achieving 82.4% VQA accuracy and 78.6% change detection precision, enabling multi-turn clinical dialogues and automated clinical reports. The framework enables dynamic, multi-turn clinical dialogues with automated medical report synthesis, addressing critical gaps in temporal reasoning, explainability, and clinical safety alignment for realworld AI and safety enhanced clinical decision support systems.
Recent advancements in multimodal learning for medical time series (MedTS) classification highlight the benefits of integrating complementary modalities for clinical decision. However, existing methods typically focus on bi-modal interactions (e.g., time series and text), leaving the tri-modal synergy between time series, vision, and language largely unexplored. Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, we introduce MedTVL, a text-guided dual-pathway architecture tailored for MedTS classification. Specifically, it synergizes a convolution-based temporal pathway for fine-grained temporal dynamics from raw numerical sequences and a transformer-based visual pathway for holistic morphological structures from time-series-derived images. Such combination of cross-modal and architectural heterogeneity provides a comprehensive diagnostic perspective. To further resolve potential diagnostic ambiguity, both pathways are guided by adaptive medical textual semantics. Finally, a Mixture-of-Experts mechanism dynamically routes each instance to specialized fusion experts, capturing instance-specific reliance on the temporal and visual pathway outputs. In addition, MedTVL supports multimodal contrastive learning to mitigate the clinical label scarcity challenge. Extensive experiments across multiple medical datasets and tasks, spanning supervised, few-shot, and contrastive learning settings, demonstrate the superiority and transferability of MedTVL, highlighting its potential for robust clinical decision support.
Jiexia Ye, Jia Li, F. Tsung· Proceedings of the 32nd ACM...· 0 citations
Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in medical scenarios and those in the general domain. Existing medical MLLMs face the following critical deficiencies: 1) inadequate coverage of medical knowledge beyond imaging; 2) elevated propensity for hallucinations due to suboptimal data curation; and 3) limited reasoning capacity tailored to complex medical tasks. To address these challenges, we first propose a comprehensive data-curation procedure that 1) efficiently acquires rich medical knowledge data not only from medical imaging but also from extensive medical texts and general domain data; and 2) synthesizes high-quality medical captions, visual question answering, and reasoning samples. Leveraging the curated data, we build a multimodal dataset imbued with extensive medical knowledge and develop our medical-specialized MLLM, Lingshu-Med, which undergoes multi-stage training to embed the medical expertise and enhance task-solving capabilities progressively. We also investigate reinforcement learning with verifiable rewards to further refine Lingshu-Med's medical reasoning abilities. For rigorous assessment, we introduce MedEvalKit, a unified evaluation framework that consolidates the leading multimodal and textual medical benchmarks for standardized, fair, and efficient model assessment. On three core medical tasks-multimodal QA, textual QA, and radiology report generation, Lingshu-Med consistently outperforms existing multimodal baselines in most tasks. Moreover, we conduct five case studies drawn from real-world clinical scenarios that illustrate its practical utility in medical contexts.
Wei-Wen Xu, H. Chan, Long Li et al.· IEEE Transactions on Pattern...· 0 citations
This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.
Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al.· 0 citations
A lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment is proposed, highlighting its efficiency and potential for clinical deployment.
Hao-Wen Gu, Gensheng Pei, Zeren Sun et al.· 2 citations
This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-of-Thought.
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text generation. Across PATH-VQA, SLAKE, and VQA-RAD, MedProb recovers substantially more answer-relevant signal than prompting and performs stronger than medical VLMs and agentic systems. Probing also reduces the apparent gap between small and large models compared to prompting, suggesting that smaller VLMs contain more recoverable Med-VQA signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve this linear decodability. Finally, free-text generation exhibits an answer-position bias of up to 10 percentage points, whereas MedProb also has positional bias, however, it is impacted differently than prompting. Our main results target the multiple-choice/multiclass Med-VQA setting; we additionally show the probe can be extended to open-ended generation via a rejection-sampling scoring procedure.
Erfan Nourbakhsh, Ke Yang, Anthony Rios· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.