Aug 2026· Journal of Medical Internet Research· Vol 28, pp. e89963-e89963· 0 citations· 43 references
Medicine
TL;DR
Evaluated large language models for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP.
Abstract
Abstract Background Large language models (LLMs) are increasingly applied in clinical decision support, yet their diagnostic performance in Chinese-language settings and under realistic clinical workflows remains unclear. In particular, how LLMs perform across diseases with different prevalence and under stepwise diagnostic processes has not been well characterized. Objective This study aimed to evaluate the diagnostic capabilities of LLMs for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework and to identify their potential and limitations for clinical diagnosis. Methods We evaluated 4 Chinese LLMs (Doubao 1.5, DeepSeek-V3, Kimi K1.5, and Leftdoctor GPT 3.5) using 56 clinical cases (28 chronic obstructive pulmonary disease [COPD], and 28 relapsing polychondritis [RP]) sourced from the China Clinical Case Results Database (March 31-April 14, 2025). Patient information was provided incrementally, starting with the initial medical history, followed by physical examination, and laboratory results. Evaluation metrics included top-3 accuracy (RTop3D), top-1 accuracy (RTopD), final diagnostic accuracy (RFA), and mean reciprocal rank (MRR). Statistical analysis was performed using generalized estimating equations (GEE), Friedman tests, and Wilcoxon signed-rank tests with Bonferroni correction. In addition, a qualitative analysis was conducted to characterize recurrent patterns of diagnostic errors. Results LLMs demonstrated significantly higher diagnostic accuracy for COPD compared to RP across all metrics (P<.001). Diagnostic accuracy improved after additional clinical information was provided, with the improvement mainly observed in RP cases. In RP, diagnostic accuracy increased from 32.14% to 71.43% for DeepSeek and from 35.71% to 78.57% for Doubao, whereas COPD accuracy remained consistently high across all diagnostic stages (82.14%‐92.86%). For COPD, ranking performance was high and comparable among all models (MRR range: 0.82‐0.89; P=.71). In RP, diagnostic performance differed significantly among models (MRR range: 0.10‐0.39; P<.001). Qualitative analysis showed that COPD errors were mainly related to a failure to recognize specific features, whereas RP errors involved more diverse patterns, particularly the neglect of negative evidence and the failure to recognize specific features. Conclusions Chinese LLMs demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP. Additional clinical information improved diagnostic accuracy primarily in RP cases, although differences between models remained evident under diagnostically complex conditions. Error patterns in RP cases suggest that current LLMs remain limited in their ability to integrate complex clinical information and exclusionary findings. Careful evaluation and appropriate clinical oversight remain important for their application in clinical practice.
The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
In the 2-stage LLM-assisted workflow, LLM assistance was associated with higher diagnostic correctness in both physician groups, although seniority-related differences in the magnitude of benefit require evaluation in larger studies.
Tusheng Li, Ziqian Ma, Baodong Wang et al.· Journal of Medical Internet...· 0 citations
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
To address the critical gap that existing medical large language model evaluation systems are predominantly based on Western medical paradigms and lack specialized assessment standards for the field of traditional Chinese medicine (TCM), this study constructs a comprehensive evaluation framework specifically designed for TCM clinical large language models, providing scientific assessment tools for their development, validation, and clinical application.
This framework addresses the critical issue of missing TCM evaluation standards by incorporating TCM-specific elements such as pattern differentiation and treatment determination, as well as formula and herb recommendations, while emphasizing patient safety and data privacy.
The framework encompasses three core dimensions: (1) clinical scenario coverage and authenticity, including renowned physician case mining, pattern identification assistance, and therapy recommendations; (2) core diagnostic and therapeutic capability support, including symptom terminology recognition, pattern identification reasoning, and medical record generation; and (3) clinical application maturity and safety assurance, including data diversity, output reliability, and privacy protection.
This evaluation system provides essential tools for the development and clinical deployment of TCM large language models, facilitating the responsible integration of AI technology into traditional medicine practice while preserving the theoretical integrity of TCM.
Nanxing Xian, Wen Zhu, Lei Zhang et al.· Guidelines and Standards of...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.