Aug 2026· SIGDIAL Conferences· pp. 821-840· 1 citation· 38 references
Computer Science
TL;DR
This work presents MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports, and introduces and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.
Abstract
Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.
This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.
Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al.· 0 citations
SUP-MIMIC is proposed, a multi-task framework utilizing MIMIC-IV-v3.1 that comprises Basic Assessment, Diagnostic Divergence Task (DDT), and Diagnostic Convergence Task (DCT), designed to evaluate the model's"one-to-many"disambiguation capability among phenotypically similar cases, while DCT assesses the model's ability to identify"many-to-one" diagnostic patterns across different pathophysiological pathways
Automated diagnostic report generation lies at the core of clinical diagnosis and can alleviate clinician shortages. Existing diagnostic report generation methods have two major limitations: they rely on unimodal inputs (e.g., images), ignoring textual biomarkers like medical history, and lack proactive dialogue capabilities to elicit personalized clinical information. To address those issues, we propose a ProActive Multimodal Agentic (PAMA) system, which performs comprehensive disease analysis by examining biomarkers in diverse sources, including multi-view medical images, medical histories, and diagnostic conversations. Built upon a knowledge graph and recommendation-based dialogue architecture, PAMA actively initiates adaptive, multi-turn conversations with patients, which is integrated with visual data for robust and reliable report generation. Specifically, PAMA actively generates adaptive multi-turn questions to collect clinically relevant background information, and then fuses the resulting dialogue context with visual representations for robust diagnostic report generation. We validate our approach on two real-world benchmark datasets, MIMIC-CXR and IU-Xray, through extensive quantitative evaluations and comparisons with state-of-the-art baselines. Furthermore, we conduct user studies to assess the realism and clinical quality of the generated reports. Finally, we present real-world case studies to examine the performance of our system across diverse scenarios, demonstrating its robustness on scalability, complexity, and data variability.
Xueshen Li, Xinlong Hou, Garrison Allen et al.· ACM Transactions on Computin...· 0 citations
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinically reliable and linguistically inclusive medical AI systems remains a significant challenge, primarily due to the lack of multimodal, multilingual, and time-series-grounded benchmarks that reflect the complexity of real-world clinical scenarios. To fill this gap, we present MMTClinic, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series. MMTClinic combines text, medical images, and multivariate physiological signals and includes 30,000 QA pairs (15,000 multiple choice questions (MCQs) and 15,000 open-ended questions) across five languages: English, Hindi, Bengali, Marathi, and Tamil. These questions cover three important clinical tasks---mortality prediction, heart rate forecasting, and SOFA score estimation. We evaluate 13 state-of-the-art LLMs in zero-shot, few-shot, and chain-of-thought settings. Our evaluation reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capabilities. MMTClinic provides a valuable resource for advancing multilingual, multimodal, and time-series-aware medical AI research. The dataset will be made publicly available on successful acceptance of the work.
Sourav Malakar, Harshit Nigam, Akash Ghosh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.