Skip to content

Author

Weihong Zheng

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study

Large language model (LLM)-based chatbots are increasingly used by the public to obtain health information, but their performance in answering questions related to traumatic brain injury (TBI) and concussion remains unclear. This study evaluated five publicly accessible LLM-based chatbots across safety, accuracy, empathy, information quality, transparency, and readability. Sixty-five English-language, public-facing questions about TBI and concussion were submitted to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao, generating 325 responses. Five blinded independent raters assessed the generated responses using guideline-informed criteria and established tools, including DISCERN, EQIP, JAMA benchmark criteria, the Global Quality Score, and readability indices. Inter-rater agreement was high. The proportions of responses rated as safe ranged from 89.2 to 95.4%, and all models achieved a median accuracy score of 4.00. However, 27 responses were classified as potentially harmful, mainly because of under-triage, overly reassuring advice regarding imaging findings, premature return-to-activity or driving guidance, and insufficient pediatric caution. Accuracy differences were statistically significant but small, whereas empathy, information quality, transparency, and readability showed clearer model-level variation. DeepSeek produced the easiest-to-read responses. LLM-based chatbots generated responses that were generally rated as safe and informative by expert evaluators, but potentially harmful advice and readability problems remained. These findings characterize expert-rated response performance and do not establish patient comprehension, educational effectiveness, or clinical benefit. LLM-based chatbots may have potential as supplementary sources of patient-facing health information, but they should not replace professional medical evaluation.

Xin Zuo, Huan Zuo, Min Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.