Skip to content
Conference

Toward Multimodal AI for Dementia: A Challenge-Driven Survey of LLMs and VLMs

Jul 2026 · International Conference on Digital Health · pp. 221-229 · 0 citations · 44 references

Abstract

Dementia affects over 57 million people worldwide and places an immense burden on informal caregivers, yet current AI tools remain largely fragmented across isolated tasks and modalities. Recent large language models (LLMs) and vision-language models (VLMs) offer promising capabilities for dementia support, but adapting them to this safety-critical, multimodal, and deeply individualized care domain raises challenges that general-purpose AI surveys do not address. In this paper, we present a challenge-driven survey that organizes the rapidly growing literature on LLMs and VLMs for dementia care around three core adaptation challenges: (1) Knowledge Grounding, which anchors model outputs to verified clinical knowledge through retrieval augmented generation, knowledge graphs, and constrained training to mitigate hallucination risk; (2) Multimodal Understanding, which fuses visual, audio, and sensor data with language to reason about the multimodal inherent of daily care; and (3) Personalization, which adapts model behavior to individual patient histories, caregiver needs, and unique disease progression over time via persistent memory and biography-driven interaction. We review over 20 recent methods, identify cross-cutting architectural patterns, and survey available datasets and benchmarks. Finally, we highlight critical open challenges including the need for standardized evaluation protocols, longitudinal deployment studies, and tighter integration between clinical workflows and foundation model capabilities.

View source

Similar papers

Sep 2026

Lingshu: Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning.

Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in medical scenarios and those in the general domain. Existing medical MLLMs face the following critical deficiencies: 1) inadequate coverage of medical knowledge beyond imaging; 2) elevated propensity for hallucinations due to suboptimal data curation; and 3) limited reasoning capacity tailored to complex medical tasks. To address these challenges, we first propose a comprehensive data-curation procedure that 1) efficiently acquires rich medical knowledge data not only from medical imaging but also from extensive medical texts and general domain data; and 2) synthesizes high-quality medical captions, visual question answering, and reasoning samples. Leveraging the curated data, we build a multimodal dataset imbued with extensive medical knowledge and develop our medical-specialized MLLM, Lingshu-Med, which undergoes multi-stage training to embed the medical expertise and enhance task-solving capabilities progressively. We also investigate reinforcement learning with verifiable rewards to further refine Lingshu-Med's medical reasoning abilities. For rigorous assessment, we introduce MedEvalKit, a unified evaluation framework that consolidates the leading multimodal and textual medical benchmarks for standardized, fair, and efficient model assessment. On three core medical tasks-multimodal QA, textual QA, and radiology report generation, Lingshu-Med consistently outperforms existing multimodal baselines in most tasks. Moreover, we conduct five case studies drawn from real-world clinical scenarios that illustrate its practical utility in medical contexts.

Wei-Wen Xu, H. Chan, Long Li et al. · 0 citations
Open access Jul 2026

A Locally Executable AI System for Improving Preoperative Patient Communication: Multidomain Clinical Evaluation

By decoupling clinical information retrieval from generative chitchat, LENOHA enhances safety, preserves privacy, and markedly reduces energy use, offering a practical blueprint for sustainable and equitable medical AI deployment across diverse care settings.

Motoki Sato, Sou Nagata, Mizuho Ohnuma et al. · 0 citations
Review Open access Jul 2026

A multimodal evidence-driven framework for clinical decision support in cognitive impairment.

The Multimodal Evidence-Driven Reasoning Framework (MEDRF), which integrates a Multimodal Hierarchical Cascade classifier with a retrieval-augmented large language model (RAG-LLM) for evidence-guided reasoning, provides a robust interpretable framework for decision support, particularly in incomplete or diagnostically ambiguous presentations.

Shicong Hu, Xiaoyang Sheng, Feng-ao Wang et al. · 0 citations
Review Open access Aug 2026

Toward Trustworthy AI for Autism Spectrum Disorder: A Systematic Review of Multimodal Systems, Knowledge Representation, and Clinical Integration

It is argued that meaningful clinical impact will require the integration of multimodal learning, semantic knowledge representation, explainable reasoning, and human-in-the-loop decision processes to support safe, interpretable, and clinically deployable AI systems in pediatric healthcare environments.

R. Zgheib, Alia El Naggar, Arash Kermani Kolankeh et al. · 0 citations
Review Aug 2026

Volumetric Radiology AI in the Era of Multimodal Large Language Models

This Review examines volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction, and introduces a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation.

Zanting Ye, Shengyuan Liu, Xin Liu et al. · 0 citations
#large language models Review Open access Sep 2026

Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings

Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resource healthcare environments where diagnostic disparities are most severe remains limited and poorly understood. This mini review synthesizes key developments in LLM–vision fusion architectures from 2018 to 2026, with a focus on radiology-oriented visual question answering (VQA) and report generation systems viewed from a deployment perspective. Rather than comprehensively cataloguing multimodal medical AI, we synthesize the evolution of LLM–vision fusion architectures and discuss complementary deployment-enabling strategies, including parameter-efficient adaptation, post-training quantization, federated learning, and multilingual support, where they directly improve the feasibility of radiology AI in resource-constrained healthcare settings. Rather than focusing solely on performance benchmarks, we examine these approaches through a deployment-oriented lens, highlighting trade-offs between representational capacity, computational efficiency, interpretability, and memory footprint. We argue that current progress remains substantially shaped by model scaling and benchmark optimization, which often do not address the memory, connectivity, and annotation constraints of low-resource healthcare systems. While cross-modal transformer architectures provide strong representational alignment, their computational demands and reliance on large curated datasets limit real-world deployment. In contrast, emerging directions including parameter-efficient fine-tuning, post-training quantization, federated learning, and modular agent-based systems offer more tractable pathways toward clinical integration under hardware and data constraints. To bridge the gap between benchmark performance and clinical utility, we identify concrete challenges in data scarcity, multilingual coverage, and calibration, and propose a shift toward lightweight, interpretable, and hardware-aware multimodal AI. This perspective highlights the need to move beyond scaling-centric design toward models that can run on 4–8 GB VRAM, operate offline, and generalize across languages and imaging equipment.

Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.