Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges
Abstract
Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed. Objectives: To determine whether diagnostic accuracy and alignment with human reader responses move together across widely accessible LLMs under default, point-of-care conditions. Methods: This cross-sectional evaluation submitted the 29 most recent New England Journal of Medicine (NEJM) case challenges to 5 publicly accessible LLMs (GPT-5, Claude, Google Gemini, DeepSeek, and Grok) through each model’s web interface (November 5-8, 2025) at default settings, one query per case, no system prompt. Each case offered a clinical vignette and 6 candidate diagnoses; published human reader response distributions (4386-17 714 per case; median, 7732) were the comparator. Diagnostic accuracy against the published gold-standard diagnosis, compared with the human reader majority by the exact McNemar test; agreement by Cohen κ; alignment by a human-error similarity index (the share of model errors on cases the reader majority also missed) and by the Jensen-Shannon divergence (JSD) between each model’s response and the reader distribution; and ordinary least squares regression of per-case JSD on normalized reader-disagreement entropy. Results: Reader majority accuracy was 48.3% (95% CI, 31.4-65.6). All 5 reached or exceeded this value, from 51.7% for DeepSeek (95% CI, 34.4-68.6) to 82.8% for Grok (95% CI, 65.5-92.4); Grok (P = .02), GPT-5 (P = .004), and Gemini (P = .04) significantly exceeded the reader majority, whereas Claude and DeepSeek did not. Grok, the most accurate model, showed the weakest alignment (similarity index, 0.40; mean JSD, 0.50 bits), with 3 of 5 errors on cases readers answered correctly. GPT-5 showed the strongest alignment (index, 1.000), failing only on cases that also challenged readers. For 4 of 5 models, divergence rose with reader-disagreement entropy; DeepSeek was the exception. Conclusions and Relevance: In this evaluation, diagnostic accuracy and alignment with human reader responses diverged across LLMs. Because a model can outperform clinicians in aggregate yet fail where oversight is least likely to detect it, evaluation of clinical AI should weigh error structure alongside accuracy.