Skip to content

Author

Mahmud Omar

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

AI-assisted rheumatology triage changes with referral framing.

OBJECTIVES Two in three US physicians now use healthcare AI, and large language models (LLMs) are entering the triage workflows that determine which patients reach rheumatology and how quickly. We aimed to test whether nine prespecified cues in referral notes, patient descriptions and demographics shift AI-assisted triage decisions when the underlying clinical information is unchanged. METHODS We conducted a controlled, physician-validated experiment across 30 physician-authored rheumatology vignettes, each independently rephrased three times (90 case variants). We tested five LLMs from three providers - Anthropic, Google, and OpenAI - under nine dimensions spanning demographics, clinical context, and communication framing, with 57 controlled contextual modifications, four system-prompt personas, and five repetitions per cell, yielding more than 200,000 model queries. Sixteen clinical fields were graded against physician-validated ground truth, with excellent inter-rater agreement (Fleiss' kappa=0.92). RESULTS Baseline composite concordance with expert ground truth was high at 0.869. We had expected sociodemographic cues to be the strongest source of distortion. Instead, the largest shifts came from how the case was framed and described. When patients were described as anxious, models attributed symptoms to psychological rather than organic causes nearly three times as often as at baseline (13.1% vs 4.5%; odds ratio 3.2 versus stoic framing), a shift that risks relabeling organic disease as functional. Clinician anchoring in the referral note reduced concordance, consistently across models and rephrasings and significantly for acuity (dismissive anchor, vignette-level p = 0.01), mainly by downgrading urgency. In contrast, race or ethnicity, socioeconomic status, and language barrier produced no detectable effect, including in mixed-effects models that accounted for repeated vignette use. CONCLUSION Although baseline concordance with specialist ground truth was high, it was readily disrupted by how referral notes were worded and how patients described their symptoms, not by patient demographics. Before AI-assisted triage enters rheumatology referral pathways, systems should separate objective clinical evidence from interpretive framing, and urgency and psychological attribution should be treated as auditable safety signals.

Mahmud Omar, Mohammad E. Naffaa, R. Agbareia et al. · 0 citations
Review Open access Jul 2026

Generative large language models in the clinical management of Alzheimer’s disease and mild cognitive impairment

Dementia affects over 55 million people worldwide. Mild cognitive impairment (MCI) often precedes Alzheimer’s disease (AD). Clinical management requires integrating uncertain evidence from neuropsychological testing, neuroimaging, and biomarkers. Large language models (LLMs) also generate probabilistic outputs, but whether they can reliably support diagnostic, therapeutic, or educational tasks in AD and MCI has not been systematically examined. We searched PubMed, Scopus, and PubMed Central (January 2023 to April 2026) for studies evaluating generative LLMs on clinical tasks in Alzheimer’s disease (AD) or mild cognitive impairment (MCI). Risk of bias was assessed using QUADAS-AI and AXIS. Narrative synthesis followed the SWiM guideline. PROSPERO: CRD420261372436. Eleven studies were included: diagnosis (n = 3), treatment guidance (n = 2), and patient/caregiver education (n = 8); two studies contributed to multiple domains. Diagnostic models achieved high internal accuracy (0.94–0.97) but declined on external validation; three-way classification accuracy dropped approximately 7% points, and MMSE-prediction R² collapsed from 0.90 to 0.25 on an external dataset. Treatment guidance approached but did not match structured clinical guidelines. Educational outputs were rated moderate to high quality but lacked source attribution and exceeded recommended reading levels; retrieval augmentation improved usability without improving accuracy. Hallucination was quantified in only 2 of 11 studies, and no study evaluated prospective clinical use. Current evidence does not support the use of LLMs for diagnosis, treatment selection, or patient education in AD/MCI without clinician oversight. These findings reflect the specific model versions, prompting strategies, and evaluation conditions in place at the time of each study, and are further limited by small heterogeneous evaluations, sparse hallucination measurement, and absence of prospective clinical validation.

Y. Adiniaev, Mahmud Omar, Oved Daniel et al. · 0 citations
Open access Aug 2026

OpenEvidence errs on the safe side in a structured test of triage recommendations.

BACKGROUND Large language model (LLM) chatbots are increasingly consulted for triage decisions. A structured benchmark reported that ChatGPT Health, a consumer-facing health assistant, under-triaged 51.6 % of true emergencies and was susceptible to social anchoring. Whether a physician-facing clinical decision support platform fails similarly is unknown. OBJECTIVE To characterize the triage safety profile of OpenEvidence, a retrieval-augmented, physician-facing platform, using the identical benchmark previously applied to ChatGPT Health. METHODS We evaluated 60 clinician-authored vignettes from 30 clinical scenarios across 21 domains. Each scenario was written with and without objective clinical data and crossed with demographic and contextual modifiers in a 2 × 2 × 2 × 2 factorial design, yielding 960 prompts (480 clear-case, 480 edge-case). Responses were classified against a clinician gold standard as correct triage, under-triage, over-triage, or evidence-seeking refusal. Analyses used cluster bootstrap resampling, mixed-effects logistic regression, and Holm-Bonferroni correction. RESULTS Among 449 clear-case responses that returned a recommendation, accuracy was 71.3 %. OpenEvidence under-triaged 12.5 % of emergency presentations versus 51.6 % in the previously reported ChatGPT Health benchmark, and over-triaged 68.0 % of nonurgent Home presentations (ChatGPT Health, 64.8 %). Anchoring statements did not alter recommendations (OR = 1.08, 95 % CI 0.62-1.88; Holm-adjusted p = 1.0). Objective clinical data eliminated emergency under-triage (25 % to 0 %; p = 0.005) and reduced nonurgent over-triage (78.7 % to 57.8 %; p = 0.014). In 65 of 960 responses (6.8 %), the platform declined to assign a triage level, exclusively in symptom-only Home or Routine prompts. CONCLUSIONS Under this benchmark, OpenEvidence produced fewer missed emergencies than the historical ChatGPT Health comparison, while errors concentrated in over-triage and evidence-seeking refusal. These findings support evaluating health AI within its deployment context and treating refusal as a distinct output category whose clinical implications require separate assessment.

Eric Jia, Mahmud Omar, Y. Barash et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.