Objectives: Large language models (LLMs) are increasingly used in medicine, but evaluation is often on multiple choice questions and management of common conditions. Infectious diseases (ID) can present complex scenarios that require considerations beyond guideline-based responses. We assessed LLM performance in these situations including with ID-specific criteria to consider infection control or antimicrobial stewardship (AMS). Methods: We evaluated four LLMs (Claude 3.5 Sonnet, GPT-4o, GPT-o1, and a local instance of Llama 3.1 8B) in October 2024, on five complex ID vignettes. The LLM responses were each evaluated for 18 items by two board-certified ID clinicians and pairwise comparisons were performed between LLMs. Results: There was no significant difference between performance of GPT-o1, GPT-4o and Claude Sonnet on general medical criteria, and were comparable with respect to how often they provided an unsafe response (GPT-o1 30%, GPT-4o 40%, Claude 37%) and contained a critical omission (GPT-o1 27%, GPT-4o 43%, Claude 47%). Llama 3.1 8B had significantly decreased performance for most criteria. On ID-specific criteria, GPT-o1 outperformed other models and all models significantly outperformed Llama for interpreting microbiology results, AMS principles, appropriate antimicrobial spectrum and infection control considerations. Performance was poor in secondary prevention and management of risk factors. Conclusions: On complex ID scenarios, LLM responses were variable. The open-source, smaller Llama 3.1 8B model performed poorly and large, non-reasoning models varied, but more than 30% of responses containing a risk of harm or critical omission. These findings suggest caution is required when deploying these models in ID domains without specialist oversight.
A. Pradhan, B. Waxse, W. Matias et al.· medRxiv· 0 citations
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem et al.· 0 citations