User Beware: Inaccuracy and Inconsistency of Large Language Models in Providing Precision Dosing Recommendations for Patients With Kidney Impairment—A Case Series
Abstract
ABSTRACT Introduction Evidence‐based dosing guidance for medications in critically ill patients with acute kidney injury (AKI) and receiving continuous kidney replacement therapy (CKRT) is limited. Freely available large language models (LLMs) can generate confident, human‐like outputs. The accuracy and reproducibility of LLMs in providing precision drug dosing recommendations in the setting of AKI and CKRT have not been evaluated. Objectives We sought to characterize literature concordance and internal consistency of LLM‐generated dosing recommendations for cefepime and meropenem in patients with AKI and receiving CKRT. Methods Six investigators queried the freely available versions of six LLMs (ChatGPT, Claude, Google Gemini, Microsoft Copilot, OpenEvidence, and Perplexity) from July to September 2025 using three standardized vignettes asking for dosing recommendations and rationale to meet prespecified pharmacodynamic targets: (i) adult with AKI not on dialysis receiving cefepime, (ii) child receiving CKRT and cefepime, and (iii) toddler receiving high‐effluent CKRT and meropenem. To assess inter‐iteration consistency, each investigator also prompted one LLM three times for each vignette. Responses were parsed for concordance with recommendations in primary literature. LLM “reasoning” was evaluated for use of pharmacokinetic (PK) equations, citation accuracy versus confabulation, acknowledgment of uncertainty, and recommendations for therapeutic drug monitoring (TDM) for efficacy or safety. Results LLM‐generated dosing recommendations varied widely. Mean concordance with literature‐based recommendations was 63% (range: 28%–94%). LLMs produced variable responses to the same user upon multiple iterations, with variance in daily maintenance doses recommended ranging from 0 to 1000%. OpenEvidence universally cited relevant sources, whereas other LLMs leveraged sources inconsistently or confabulated them. Each recommended TDM, though only Claude consistently acknowledged its own uncertainty. Conclusion Freely available LLMs produce highly variable and often discordant antibiotic dosing recommendations for patients with AKI and receiving CKRT. Although valuable for hypothesis generation and literature retrieval, LLM outputs should not be used in isolation for drug dosing in critically ill patients with kidney dysfunction.