Skip to content
Open access

Same child, different risk: demographic bias in childhood obesity attribution by large language models

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 36 references
Medicine

TL;DR

Publicly accessible English-language web-interface outputs from current LLMs showed systematic demographic patterns in pediatric obesity risk attribution, supporting the need for pre-deployment and post-deployment bias auditing before clinical or consumer health use.

Abstract

Background Large language models (LLMs) are increasingly consulted for pediatric health information, yet their demographic biases remain unsystematically evaluated in pediatric contexts. Objectives To assess bias and variability in childhood obesity risk attribution across seven LLMs (ChatGPT, Claude, DeepSeek, Gemini, GLM, Grok, and Qwen), spanning both Western and Chinese-origin developers; all prompts, including those submitted to the Chinese-origin models, were in English only. Methods A structured prompt-based experimental design was employed across six clinical domains (general obesity risk, dietary pattern, physical activity, sleep, mental health, and genetic predisposition) and six demographic comparison dimensions (sex, three race/ethnicity pairings, socioeconomic status, and urban-rural residence). Seventy-eight unique prompts were submitted to each model in triplicate, yielding 1,638 outputs. Neutral prompts were scored on a five-dimension binary rubric (accuracy, representation, stigmatizing/harmful language, social determinants, cultural fit); comparative prompts were coded for directional risk attribution. Results Claude achieved the highest neutral prompt composite score (mean 3.00 ± 0.91) and GLM the lowest (1.44 ± 0.51); between-model differences were statistically significant (Kruskal–Wallis H = 46.21, p < 0.001). All models achieved a 100% Stigmatizing/Harmful Language pass rate, yet representation and cultural fit were universally weak. Socioeconomic status produced the most consistent attribution pattern (low-income attribution in 40/42 decisions; decision change rate 19.0%). Most models attributed higher obesity risk to Black and Hispanic/Latino children across the majority of domains. Urban–rural attribution showed the greatest cross-model directional inconsistency (decision change rate 52.4%), with Western-origin models favoring rural attribution and Chinese-origin models favoring urban attribution. Conclusions Publicly accessible English-language web-interface outputs from current LLMs showed systematic demographic patterns in pediatric obesity risk attribution, supporting the need for pre-deployment and post-deployment bias auditing before clinical or consumer health use.

Read PDF

Similar papers

#large language models Open access Sep 2026

GPT-4 improves sex-specificity in cardiovascular patient education but may perpetuate gender biases: A mixed-methods audit

Patient education materials for cardiovascular disease (CVD) prevention frequently omit clinically important differences in disease manifestation, risk, and prevention between sexes. Socially constructed gender norms further shape how health information is communicated and received. Large Language Models (LLMs) like GP...

Samah Khan, G. Vaidean · 0 citations
Open access Aug 2026

Poster 162. Racial Biases Perpetuated by Modern Large Language Models Negatively Impact Diagnostic Reasoning and Treatment Recommendations in Musculoskeletal Healthcare

Objectives: To determine whether contemporary large language models (LLMs) perpetuate gender and racial biases in medical decision-making. Methods: A total of 180 standardized vignettes concerning musculoskeletal diagnoses and treatments were extracted from AAOS Restudy and Orthobullets. A customized GPT-o4-mini model...

Kyle N. Kunze, Nicholas Allen, Sophia J. Madjarova et al. · 0 citations
Review Open access Aug 2026

Synthesizing Risk Factors for Alcohol Use Disorder Using a Large Language Model

These findings demonstrate the value of an artificial intelligence-driven literature review for informing comprehensive strategies to ad- dress the multifactorial nature of AUD.

Chen-Lan Wang, Kurtis Riener, Yue LuoSan · 0 citations
Open access Sep 2026

Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis

Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance...

Li Xu, Xu Qiu, Jia-Yi Deng et al. · 0 citations
Open access Aug 2026

Digital twin–supported behavioral intention in mothers of young children to prevent childhood obesity: a large language model–based intervention study

Introduction Childhood obesity is a major determinant of lifelong non-communicable disease risk, yet early intervention may modify this trajectory. Digital twins and large language models may provide a scalable means of translating individualized risk prediction into understandable and motivating lifestyle guidance. Me...

Kenji Nakamura, G. Shinoda, A. Noda et al. · 0 citations
Review

World Journal of Clinical Pediatrics

Hugo Estupiñán-Pérez, Angela M. Jiménez-Urrego, A. Carvajal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.