Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

OpenEvidence errs on the safe side in a structured test of triage recommendations.

BACKGROUND Large language model (LLM) chatbots are increasingly consulted for triage decisions. A structured benchmark reported that ChatGPT Health, a consumer-facing health assistant, under-triaged 51.6 % of true emergencies and was susceptible to social anchoring. Whether a physician-facing clinical decision support platform fails similarly is unknown. OBJECTIVE To characterize the triage safety profile of OpenEvidence, a retrieval-augmented, physician-facing platform, using the identical benchmark previously applied to ChatGPT Health. METHODS We evaluated 60 clinician-authored vignettes from 30 clinical scenarios across 21 domains. Each scenario was written with and without objective clinical data and crossed with demographic and contextual modifiers in a 2 × 2 × 2 × 2 factorial design, yielding 960 prompts (480 clear-case, 480 edge-case). Responses were classified against a clinician gold standard as correct triage, under-triage, over-triage, or evidence-seeking refusal. Analyses used cluster bootstrap resampling, mixed-effects logistic regression, and Holm-Bonferroni correction. RESULTS Among 449 clear-case responses that returned a recommendation, accuracy was 71.3 %. OpenEvidence under-triaged 12.5 % of emergency presentations versus 51.6 % in the previously reported ChatGPT Health benchmark, and over-triaged 68.0 % of nonurgent Home presentations (ChatGPT Health, 64.8 %). Anchoring statements did not alter recommendations (OR = 1.08, 95 % CI 0.62-1.88; Holm-adjusted p = 1.0). Objective clinical data eliminated emergency under-triage (25 % to 0 %; p = 0.005) and reduced nonurgent over-triage (78.7 % to 57.8 %; p = 0.014). In 65 of 960 responses (6.8 %), the platform declined to assign a triage level, exclusively in symptom-only Home or Routine prompts. CONCLUSIONS Under this benchmark, OpenEvidence produced fewer missed emergencies than the historical ChatGPT Health comparison, while errors concentrated in over-triage and evidence-seeking refusal. These findings support evaluating health AI within its deployment context and treating refusal as a distinct output category whose clinical implications require separate assessment.

Eric Jia, Mahmud Omar, Y. Barash et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.