Dementia affects over 55 million people worldwide. Mild cognitive impairment (MCI) often precedes Alzheimer’s disease (AD). Clinical management requires integrating uncertain evidence from neuropsychological testing, neuroimaging, and biomarkers. Large language models (LLMs) also generate probabilistic outputs, but whether they can reliably support diagnostic, therapeutic, or educational tasks in AD and MCI has not been systematically examined. We searched PubMed, Scopus, and PubMed Central (January 2023 to April 2026) for studies evaluating generative LLMs on clinical tasks in Alzheimer’s disease (AD) or mild cognitive impairment (MCI). Risk of bias was assessed using QUADAS-AI and AXIS. Narrative synthesis followed the SWiM guideline. PROSPERO: CRD420261372436. Eleven studies were included: diagnosis (n = 3), treatment guidance (n = 2), and patient/caregiver education (n = 8); two studies contributed to multiple domains. Diagnostic models achieved high internal accuracy (0.94–0.97) but declined on external validation; three-way classification accuracy dropped approximately 7% points, and MMSE-prediction R² collapsed from 0.90 to 0.25 on an external dataset. Treatment guidance approached but did not match structured clinical guidelines. Educational outputs were rated moderate to high quality but lacked source attribution and exceeded recommended reading levels; retrieval augmentation improved usability without improving accuracy. Hallucination was quantified in only 2 of 11 studies, and no study evaluated prospective clinical use. Current evidence does not support the use of LLMs for diagnosis, treatment selection, or patient education in AD/MCI without clinician oversight. These findings reflect the specific model versions, prompting strategies, and evaluation conditions in place at the time of each study, and are further limited by small heterogeneous evaluations, sparse hallucination measurement, and absence of prospective clinical validation.
Y. Adiniaev, Mahmud Omar, Oved Daniel et al.· Neurological Sciences· 0 citations
BACKGROUND
Large language model (LLM) chatbots are increasingly consulted for triage decisions. A structured benchmark reported that ChatGPT Health, a consumer-facing health assistant, under-triaged 51.6 % of true emergencies and was susceptible to social anchoring. Whether a physician-facing clinical decision support platform fails similarly is unknown.
OBJECTIVE
To characterize the triage safety profile of OpenEvidence, a retrieval-augmented, physician-facing platform, using the identical benchmark previously applied to ChatGPT Health.
METHODS
We evaluated 60 clinician-authored vignettes from 30 clinical scenarios across 21 domains. Each scenario was written with and without objective clinical data and crossed with demographic and contextual modifiers in a 2 × 2 × 2 × 2 factorial design, yielding 960 prompts (480 clear-case, 480 edge-case). Responses were classified against a clinician gold standard as correct triage, under-triage, over-triage, or evidence-seeking refusal. Analyses used cluster bootstrap resampling, mixed-effects logistic regression, and Holm-Bonferroni correction.
RESULTS
Among 449 clear-case responses that returned a recommendation, accuracy was 71.3 %. OpenEvidence under-triaged 12.5 % of emergency presentations versus 51.6 % in the previously reported ChatGPT Health benchmark, and over-triaged 68.0 % of nonurgent Home presentations (ChatGPT Health, 64.8 %). Anchoring statements did not alter recommendations (OR = 1.08, 95 % CI 0.62-1.88; Holm-adjusted p = 1.0). Objective clinical data eliminated emergency under-triage (25 % to 0 %; p = 0.005) and reduced nonurgent over-triage (78.7 % to 57.8 %; p = 0.014). In 65 of 960 responses (6.8 %), the platform declined to assign a triage level, exclusively in symptom-only Home or Routine prompts.
CONCLUSIONS
Under this benchmark, OpenEvidence produced fewer missed emergencies than the historical ChatGPT Health comparison, while errors concentrated in over-triage and evidence-seeking refusal. These findings support evaluating health AI within its deployment context and treating refusal as a distinct output category whose clinical implications require separate assessment.
Eric Jia, Mahmud Omar, Y. Barash et al.· International Journal of Med...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.