Jun 2026· arXiv.org· Vol abs/2606.31608· 0 citations· 33 references
Computer Science
Abstract
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce CLExEval, a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking. CLExEval combines 5,600 expert-physician annotations with 200 clinical reasoning traces derived from 40 rare diagnostic cases. Our analysis identifies three recurring failure patterns: (i) verbosity bias, where GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity; (ii) a hidden knowledge paradox, where a specialist model reaches 92.5% maximum diagnostic potential but fails to retrieve that knowledge reliably in verbose contexts; and (iii) a 68.6% reasoning-to-output mismatch, where correct diagnoses appear in reasoning traces but are not reflected in final answers. We further evaluate the LLM-as-a-Judge paradigm on a human-verified failure set (n = 142). GPT-4o-mini approved 47.9% of clinically incorrect outputs, while HuatuoGPT-o1 approved all validly scored failures and showed a positive self-preference bias. These results suggest that standalone automated clinical evaluations can substantially overestimate clinical reliability without expert-grounded validation.
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable clinically-grounded reasoning from pattern matching. We introduce clinical reasoning graphs, structured graph representations extracted from free-text LLM diagnostic traces using a domain-grounded ontology with 5 node types and 7 edge types. We apply this pipeline to 750 traces from five LLMs across 50 New England Journal of Medicine Clinicopathological Conference cases and three prompt conditions, and test whether diagnostic traces show stable structured reasoning patterns, or diagnostic schemas, for clinically similar cases. We operationalize this as higher graph similarity among clinically similar cases than among clinically dissimilar ones. Across 15 model-condition comparisons, within-cluster and between-cluster composite similarity are nearly equal, and no comparison survives multiple-testing correction; a component-level analysis finds any residual content signal far below schema scale. Graph similarity is also nearly identical for pairs of models that are both correct (0.488) and both incorrect (0.484), suggesting that graph structure captures a dimension not reflected in diagnostic accuracy. Structured reflection prompting increases explicit discriminating-feature analysis within traces (+33%) but does not increase cross-case consistency. These results show diagnostic competence without schema-scale reasoning consistency, and indicate that final-answer accuracy should be complemented by process-level evaluation. We release the ontology, extraction pipeline, validation protocol, and the extracted reasoning graphs and similarity artifacts as resources for structured evaluation of LLM clinical reasoning.
Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematologic oncology in which models must proactively request clinical data across three sequential rounds before committing to a diagnosis and treatment plan. Across 32 frontier models, the best achieved only 68% overall accuracy. Information utilization, the fraction of available data actually requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P<0.001), yet utilization collapsed from 57% to 26% in the final round, leaving molecular and cytogenetic data critical for treatment selection unexamined. Reasoning traces scored high on a clinical reasoning rubric (91% above threshold) but decorrelated from accuracy, revealing a gap between locally coherent rationales and globally correct conclusions. Error analysis identified search satisficing, anchoring and premature closure as the dominant failure modes, the same cognitive biases that characterize novice clinicians under dual-process models of diagnostic reasoning. These findings demonstrate that the primary limitation of current models in clinical oncology is not insufficient medical knowledge but a systematic failure of information-seeking under uncertainty.
K. Braitsch, L. Schmalbrock, Theresa Weltermann et al.· 1 citation
Large language model (LLM)-as-a-judge systems offer scalable evaluation of artificial intelligence (AI)-generated clinical outputs, yet their susceptibility to prompt variability raises concerns regarding reproducibility and alignment with expert judgement. This study examined whether evaluation prompt strategies influence scoring patterns and concordance with clinical raters in critical care.
This post-hoc analysis used 90 structured clinical reports generated in a prior study using an XGBoost ICU mortality prediction model trained on the MIMIC-IV database. GPT-4o (Azure AI, version 2024-11-20) produced structured interpretations from risk estimates and SHAP attributions. These outputs were evaluated using the IMPACT framework under three evaluation prompt strategies: baseline (E1), top-down decremental (E2), and bottom-up incremental (E3). Agreement between clinician ratings and the automated o3-mini evaluator (Azure AI, version 2025-01-31) was assessed using intraclass correlation coefficients (ICC), with strategy comparisons by Fisher’s z-transformation. Score deviations were examined with repeated-measures ANOVA.
Mean IMPACT scores were 79.9 (SD 9.9) for E1, 83.3 (SD 9.6) for E2, 78.7 (SD 9.1) for E3, and 78.6 (SD 8.9) for clinicians. All strategies demonstrated substantial agreement (ICC > 0.80). E2 showed significantly lower agreement with clinicians (ICC = 0.82) than E1 and E3 (both ICC = 0.94, p < 0.001). Score deviations differed significantly across strategies (p < 0.001), with E3 showing the smallest mean deviation (0.1) and E2 the largest (4.7).
Prompt design meaningfully affects both IMPACT scoring patterns and the reliability of LLM-based evaluators. Bottom-up incremental scoring showed the closest alignment with human assessment, underscoring the need for standardised prompt architectures in clinical AI evaluation.
Stakeholder role-prompting fundamentally alters clinical decisions and ethical value frameworks of frontier LLMs, with the insurer role producing systematic denial of physician-endorsed, patient-preferred treatments.
C. Dave, A. Diviero, T. Dassanayake et al.· medRxiv· 0 citations