Skip to content

Author

Zhao-Le Gong

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Evaluation of generative AI-driven chatbots as sources of consumer health information on hand, foot, and mouth disease: a cross-sectional comparative study of safety, accuracy, information quality, readability, and empathy

Background Generative artificial intelligence chatbots are increasingly used as sources of consumer health information. Although their performance has been examined in several medical conditions, evidence specific to hand, foot, and mouth disease (HFMD) remains limited. Objective To compare the safety, accuracy, empathy, information quality, and readability of HFMD-related responses generated by five publicly accessible chatbots. Methods In this exploratory cross-sectional study, 20 researcher-developed English-language prompts were constructed from authoritative public-health sources, Google Trends topic mapping, and caregiver-informed wording refinement. ChatGPT-4o, Gemini 2.5 Pro, Copilot, Doubao, and DeepSeek-V3.2-Exp were evaluated between April 2 and April 5, 2026. Five trained reviewers independently assessed responses against predefined reference standards using safety and accuracy criteria, an empathy scale, DISCERN, EQIP, the Global Quality Score, and JAMA benchmarks. Six formula-based readability indices were calculated. Paired comparisons used Friedman and Cochran Q tests, with prespecified post-hoc procedures and Benjamini-Hochberg correction. Results Unsafe-response rates ranged from 5.0 to 15.0%, with no statistically significant difference detected among chatbots (p = 0.797). No statistically significant inter-model differences were detected for accuracy, empathy, DISCERN, EQIP, JAMA, or Global Quality Score in the 20-prompt set; because the study was exploratory and was not powered to establish equivalence, these findings do not demonstrate comparable or interchangeable performance. All six readability indices differed significantly among chatbots (p = 0.030 to <0.001). ChatGPT and Doubao generally produced lower estimated grade-level complexity than Gemini and DeepSeek. Ten potentially unsafe or misleading responses were identified, mainly involving overgeneralization of EV71 vaccine protection, hand-hygiene qualification, and disinfection advice. Conclusion Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained. These systems may support general information seeking, but their responses require cautious interpretation and should not replace individualized professional advice.

Zhao-Le Gong, Yan Na, Yi Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.