A machine-assisted framework for systematic error analysis in clinical concept extraction
Abstract
Error analysis is critical for evaluating and improving clinical natural language processing (NLP) models, yet no standardized framework exists to support systematic error analysis for clinical concept extraction. Here we present MedError, a machine-assisted, human-in-the-loop framework integrating a validated error taxonomy, large language model (LLM)-assisted classification, and a structured annotation workflow. MedError is developed from 1187 unique errors manually curated from 4227 clinical notes across three institutions, spanning 25 error types and 48 clinical concept categories. Here we show that MedError supports both single-site and federated multisite error analysis through three complementary evaluations. First, benchmarking of six proprietary and open-source LLMs shows that automated error classification alone is insufficient without human involvement. Second, MedError-derived error feedback improves NLP extraction performance by 34–62% over an unstructured baseline across 11 cognitive status concepts. Third, prospective external validation at two independent institutions demonstrates improved agreement with expert-adjudicated annotations by up to 30% (McNemar exact test, p = 0.035, φ = 0.38) and reduced annotation time by 26–33%, providing a systematic and practical framework for real-world clinical NLP deployment. MedError, a machine-assisted human-in-the-loop framework, improves the standardization, validity, and efficiency of clinical NLP error analysis, supporting more transparent and trustworthy model deployment.