Skip to content

Author

Benedict U. Nwachukwu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Poster 162. Racial Biases Perpetuated by Modern Large Language Models Negatively Impact Diagnostic Reasoning and Treatment Recommendations in Musculoskeletal Healthcare

Objectives: To determine whether contemporary large language models (LLMs) perpetuate gender and racial biases in medical decision-making. Methods: A total of 180 standardized vignettes concerning musculoskeletal diagnoses and treatments were extracted from AAOS Restudy and Orthobullets. A customized GPT-o4-mini model was prompted using designs that resembled typical use within clinical settings and fed vignettes via a custom queryLLM2 pipeline, programmatically generating JSON-formatted differentials across demographic groups (Caucasian/African-American/Asian/Hispanic and Male/Female) while holding all other content constant. JSON outputs were parsed to extract ordered diagnosis lists, compute the rank of the reference ('correct') diagnosis and a top-3 inclusion indicator, and derive sentiment scores. Differential diagnosis and treatment planning were evaluated across demographic groups using Kruskal-Wallis tests and pairwise Mann-Whitney U tests with Benjamini-Hochberg FDR correction, alongside standardized proportion and rank-delta visualizations. Results: Among 1,440 race/gender combinations, the model demonstrated outputs that were more likely to recommend diagnoses and treatments that stereotyped certain racial groups (Figure 1). Furthermore, the model was significantly more likely to provide the correct diagnosis for Asian and Caucasian patients, while the proportion of correct diagnoses within the top 3 diagnoses listed was significantly lower for African American and Hispanic patients (56% vs. 27%, p<0.05). Gender was not significantly associated with different diagnostic or treatment rankings. Conclusions: Contemporary LLMs may perpetuate racial biases acquired during model training when being used to reason through musculoskeletal healthcare content. These findings highlight a concerning limitation in the use of LLMs and therefore there is a need for enhanced transparency and mitigation of these biases prior to integration into clinical workflows.

Kyle N. Kunze, Nicholas Allen, Sophia J. Madjarova et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.