Skip to content

MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

Aug 2026 · SIGDIAL Conferences · pp. 821-840 · 1 citation · 38 references
Computer Science

TL;DR

This work presents MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports, and introduces and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.

Abstract

Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.

View source

Similar papers

Preprint Aug 2026

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.

Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs'Robustness to Contradictory Evidence

SUP-MIMIC is proposed, a multi-task framework utilizing MIMIC-IV-v3.1 that comprises Basic Assessment, Diagnostic Divergence Task (DDT), and Diagnostic Convergence Task (DCT), designed to evaluate the model's"one-to-many"disambiguation capability among phenotypically similar cases, while DCT assesses the model's ability to identify"many-to-one" diagnostic patterns across different pathophysiological pathways

Yi Yu, Bo Wang, Chong Feng et al. · 0 citations
Jul 2026

Diagnostic Report Generation via a ProActive Multimodal Agentic System

Automated diagnostic report generation lies at the core of clinical diagnosis and can alleviate clinician shortages. Existing diagnostic report generation methods have two major limitations: they rely on unimodal inputs (e.g., images), ignoring textual biomarkers like medical history, and lack proactive dialogue capabilities to elicit personalized clinical information. To address those issues, we propose a ProActive Multimodal Agentic (PAMA) system, which performs comprehensive disease analysis by examining biomarkers in diverse sources, including multi-view medical images, medical histories, and diagnostic conversations. Built upon a knowledge graph and recommendation-based dialogue architecture, PAMA actively initiates adaptive, multi-turn conversations with patients, which is integrated with visual data for robust and reliable report generation. Specifically, PAMA actively generates adaptive multi-turn questions to collect clinically relevant background information, and then fuses the resulting dialogue context with visual representations for robust diagnostic report generation. We validate our approach on two real-world benchmark datasets, MIMIC-CXR and IU-Xray, through extensive quantitative evaluations and comparisons with state-of-the-art baselines. Furthermore, we conduct user studies to assess the realism and clinical quality of the generated reports. Finally, we present real-world case studies to examine the performance of our system across diverse scenarios, demonstrating its robustness on scalability, complexity, and data variability.

Xueshen Li, Xinlong Hou, Garrison Allen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinically reliable and linguistically inclusive medical AI systems remains a significant challenge, primarily due to the lack of multimodal, multilingual, and time-series-grounded benchmarks that reflect the complexity of real-world clinical scenarios. To fill this gap, we present MMTClinic, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series. MMTClinic combines text, medical images, and multivariate physiological signals and includes 30,000 QA pairs (15,000 multiple choice questions (MCQs) and 15,000 open-ended questions) across five languages: English, Hindi, Bengali, Marathi, and Tamil. These questions cover three important clinical tasks---mortality prediction, heart rate forecasting, and SOFA score estimation. We evaluate 13 state-of-the-art LLMs in zero-shot, few-shot, and chain-of-thought settings. Our evaluation reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capabilities. MMTClinic provides a valuable resource for advancing multilingual, multimodal, and time-series-aware medical AI research. The dataset will be made publicly available on successful acceptance of the work.

Sourav Malakar, Harshit Nigam, Akash Ghosh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.