Skip to content
Preprint

Toward Better Assessment of LLMs'Performance in Clinical Error Detection

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

While models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart and it is shown that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators.

Abstract

Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart. Aggregate discriminative metrics (e.g., balanced accuracy or F1) do not exploit this structure. We show that this omission is consequential. In particular, evaluating 15 diverse LLMs on 4 standardized clinical error-detection test sets across 3 languages, we find that 13 of 15 models fall below the level of random pairwise discrimination, even while achieving F1 scores that standard practice would read as moderate. We also observe that the underlying bias patterns differ across languages: the same model can default to"no error"on one language and over-flag errors on another. To diagnose where discrimination breaks down, we further introduce a procedure to score the evidence models cite in their outputs. We find that while models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart. Finally, we show that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators. For safety-critical clinical NLP applications, we advocate for supplementing aggregate metrics with paired evaluations in benchmark reporting. Code and analysis scripts are available at https://github.com/healthylaife/paired-clinical-eval.

View source

Similar papers

Open access Jul 2026

Reducing overconfident errors in clinical prediction models.

Machine Learning (ML) models are increasingly being used in clinical workflows. Evaluation of these models tends to focus on global performance metrics, which can obscure error patterns and lead to bias in clinical decision making. Here, we propose Proximal Error-Based Confidence Adjustment (PECA), a framework designed explicitly to improve the safety of ML predictions by reducing a model's confidence in regions of the feature space associated with historical prediction errors. Rather than optimizing global metrics alone, PECA targets confident misclassifications by attenuating prediction confidence proportional to similarity with previously observed errors. Across extensive simulations, PECA reduced the rate at which a model made confident misclassifications while preserving the overall predictive ability of the model. We applied this framework to a clinical trial enrollment workflow for Alzheimer's disease and demonstrated consistent results with the simulations while also demonstrating superior statistical power compared to baseline models. The results of this paper suggest that taming a model's tendency to predict overconfidently using historical error patterns may be a critical step towards safer and more reliable ML systems for digital public health.

Henry Bayly, Y. Tripodis, Steven Lenio et al. · 0 citations
Conference Open access Jul 2026

When Confidence Fails: Overconfidence in LLMS Under Uncertainty and Missing Clinical Information

An evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings is proposed, which introduces linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts and observes significant variation across models in their ability to abstain when the correct answer is unavailable.

Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam et al. · 0 citations
Review Open access Jul 2026

A Modular Evaluation of AI-Assisted Clinical Documentation

The results support the use of this modular AI-assisted clinical documentation pipeline as a human-supervised draft-generation tool that still requires clinician review, local workflow evaluation, and prospective clinical validation before broader deployment.

Julien Delaunay, Maissaa Sarkis, Jordi Solé-Casals et al. · 0 citations
Book Open access Feb 2026

LiveMedBench: A Contamination-Limited Medical Benchmark for LLMs with Automated Rubric Evaluation

A Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles and develops an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge.

Zhiling Yan, D. Song, Zhen Fang et al. · 8 citations · ⚡2
Review Aug 2026

AI-Powered Prescription Error Detection Using Large Language Models (LLMs): A Systematic Review and Future Perspectives

It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.

K. K. Kumar, Koyya Gowtham Reddy, K. Reddy · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.