Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Sep 2026

A primer on evaluation methods for large language models in healthcare

Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.

S. E. McKinney, P. Vu, S. Justice et al. · 0 citations
Review Open access Aug 2026

An Electronic Health Record–Integrated, Large Language Model–Powered Tool to Triage Surgical Patients

Key Points Question Can surgical patient triage be automated using a large language model (LLM) agentic workflow? Findings In this quality improvement study, the LLM tool recommended hospitalist consultation for nearly a quarter of the 6193 triaged cases. The tool achieved 94% sensitivity and 74% specificity, and post hoc medical record review suggested that most discrepancies reflected modifiable gaps in clinical criteria, institutional workflow, or physician practice variability, rather than LLM misclassification. Meaning The findings of this study suggest that an LLM-powered human-in-the-loop agentic workflow could accurately triage surgical patients for a surgical comanagement service.

Janelle B. Wang, T. Keyes, April S. Liang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.