Skip to content

Author

Xiaoqing Si

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Toward safer digital sexual health communication: evaluating the public health reliability of large language model responses on sexually transmitted infections

The global burden of sexually transmitted infections (STIs) continues to rise, yet stigma drives many to seek sensitive health information from artificial intelligence chatbots. The quality, safety, readability, and destigmatization of large language model (LLM) responses on sexual health remain under-evaluated, particularly for non-Western models. This cross-sectional study constructed 30 standardized English queries (5 themes × 6 questions) informed by Google Trends, Centers for Disease Control and Prevention (CDC) guidelines, and patient education platforms, and submitted them to three LLMs—GPT-5.4, DeepSeek-V4-Pro, and Kimi K2.6—via official APIs, in duplicate (180 responses). Two dermatovenereologists (>10 years' experience) rated the responses while blinded to platform identity, across six dimensions (accuracy, completeness, safety, understandability, actionability, and destigmatization) using a five-point scale, supplemented by four readability metrics (FKGL, FRE, GFI, and SMOG). Primary analysis used linear mixed-effects models retaining all individual ratings; the original aggregated non-parametric pipeline was retained as a sensitivity analysis. Reliability used ICC and quadratic-weighted Cohen's κ with 95% CIs; correlations used query-level cluster bootstrap with false-discovery-rate control. Inter-platform differences were significant for accuracy (χ 2 (2) = 19.12, Holm-adjusted P < 0.001), understandability (χ 2 (2) = 13.54, P = 0.006), and destigmatization (χ 2 (2) = 9.16, P = 0.041). GPT-5.4 and Kimi K2.6 achieved the highest accuracy (both median 5.00), both significantly above DeepSeek-V4-Pro and not significantly different from each other. DeepSeek-V4-Pro and Kimi K2.6 both used significantly more destigmatizing language than GPT-5.4, with no significant difference between them. GPT-5.4 produced the least readable text on all four metrics (all adjusted P < 0.01). Inter-rater reliability was good to excellent (ICC(2, 1) = 0.79–0.95; quadratic-weighted κ = 0.79–0.95). The three LLMs showed distinct profiles: GPT-5.4 and Kimi K2.6 achieved the highest accuracy, whereas GPT-5.4 produced the least readable outputs; DeepSeek-V4-Pro and Kimi K2.6 used more destigmatizing language. To our knowledge, this is the first study to quantify destigmatizing language as an explicit evaluation dimension for LLM-generated STI content, albeit as an exploratory measure pending formal content validation. Findings are specific to these models and queries and can inform quality standards and monitoring for safer AI-powered sexual health communication.

Shucheng Zhang, Shuo Wang, Xiao-Yue Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.