Skip to content
Open access

Evaluating the accuracy, reliability, and readability of AI chatbots in delivering postpartum depression information

Aug 2026 · Frontiers in Psychiatry · Vol 17 · 0 citations · 30 references
Medicine

TL;DR

AI chatbots, particularly ChatGPT-5, showed potential as supplementary tools for providing postpartum depression related information, especially in standardized MCQ based assessment, but inadequate transparency, limited source attribution, and suboptimal readability indicate that they should not be used as autonomous sources of postpartum mental health guidance and should not replace professional assessment or care.

Abstract

Background Postpartum depression (PPD) is a common perinatal psychiatric disorder with significant implications for maternal and infant health. Artificial intelligence (AI) chatbots have emerged as widely accessible tools for health information, but their performance in providing accurate, reliable, and readable PPD related information remains underexplored. Methods We evaluated six AI chatbots, ChatGPT-5, ChatGPT-4o, Claude Sonnet 4.5, DeepSeek-V3.2, DeepSeek-R1, and Gemini 2.5 Pro, using 200 standardized multiple choice questions (MCQs) on PPD to assess validity. ChatGPT-4o and DeepSeek-R1 were included only in the MCQ based validity analysis as earlier version comparators. Reliability and readability were further assessed using 20 core public education questions in the four latest models: ChatGPT-5, Claude Sonnet 4.5, DeepSeek-V3.2, and Gemini 2.5 Pro. Chatbot performance was assessed across three dimensions: validity, reliability, and readability. Each MCQ was presented three times independently, and each core public education question was assessed once per model. Results In the six model MCQ based validity analysis, ChatGPT-5 achieved the highest overall accuracy on MCQs (97.50% ± 0.50%). In the four models reliability and readability analyses, ChatGPT-5 obtained the highest DISCERN, EQIP, and GQS scores, suggesting relatively better content quality and user oriented usefulness. However, JAMA benchmark scores were low across all models, including ChatGPT-5, indicating limited transparency, source attribution, currency, and disclosure. All models produced outputs exceeding the recommended sixth grade reading level, although ChatGPT-5 and Gemini 2.5 Pro were relatively more accessible. Conclusion AI chatbots, particularly ChatGPT-5, showed potential as supplementary tools for providing postpartum depression related information, especially in standardized MCQ based assessment. However, this study did not evaluate clinical safety, patient comprehension, user behavior, or real world effectiveness, and suboptimal readability may limit accessibility for users with lower health or digital literacy. Inadequate transparency, limited source attribution, and suboptimal readability indicate that AI chatbots should not be used as autonomous sources of postpartum mental health guidance and should not replace professional assessment or care.

Read PDF

Similar papers

Open access Aug 2026

Comparative Evaluation of ChatGPT-5.2, Claude Sonnet 4.5, and DeepSeek-V3.2 for Rosacea-Related Information: Accuracy, Reliability, Readability, and Reference Hallucinations

Background/Objectives: Rosacea is a chronic inflammatory skin disease that requires long-term management and continuous patient education regarding triggers, skincare practices, and treatment adherence. In recent years, patients have increasingly turned to online platforms and artificial intelligence (AI)-based chatbots for health-related information. Although ChatGPT has been evaluated in the context of rosacea, evidence regarding the performance of other AI chatbots remains limited. This study aimed to evaluate the accuracy, reliability, quality, readability, and diagnostic relevance of AI-generated responses to common rosacea-related patient questions and to assess their potential role as sources of health-related information. Methods: Between 21 December 2025 and 22 February 2026, rosacea-related questions were collected from the publicly accessible Quora platform using a systematic screening process. Twenty clinically relevant and representative questions covering diagnosis, triggers, treatment options, skincare practices, and disease manifestations were selected. Each question was independently submitted to three AI chatbots (Claude Sonnet 4.5, ChatGPT-5.2, and DeepSeek-V3.2). Responses were evaluated by domain experts using the modified DISCERN (mDISCERN) for reliability, the Global Quality Scale (GQS) for overall quality, the Flesch Reading Ease Score (FRES) for readability, and a 5-point Likert scale for accuracy. Reference hallucinations were assessed through manual verification of cited sources. Statistical comparisons were performed using the Friedman test with Bonferroni-adjusted post hoc analyses, and effect sizes were calculated using Kendall’s coefficient of concordance (Kendall’s W). Results: Significant differences were observed among the AI chatbots across all evaluation domains (p < 0.05), with moderate to large effect sizes (Kendall’s W = 0.272–0.683). ChatGPT-5.2 and DeepSeek-V3.2 demonstrated significantly higher reliability and accuracy scores than Claude Sonnet 4.5. DeepSeek-V3.2 achieved the highest overall quality scores, whereas ChatGPT-5.2 produced the most readable responses. Reference analysis revealed variable hallucination rates among the evaluated models. Conclusions: Generative AI chatbots demonstrate considerable potential as sources of health-related information for rosacea-related queries. However, variability in performance and reference hallucination rates highlights the need for careful validation before their widespread use as complementary sources of patient health information.

Mahmut Talha Uçar, Ecem Bostan, Tulay Ortabag et al. · 0 citations
Review Open access Aug 2026

Evaluation of generative AI-driven chatbots as sources of consumer health information on hand, foot, and mouth disease: a cross-sectional comparative study of safety, accuracy, information quality, readability, and empathy

Background Generative artificial intelligence chatbots are increasingly used as sources of consumer health information. Although their performance has been examined in several medical conditions, evidence specific to hand, foot, and mouth disease (HFMD) remains limited. Objective To compare the safety, accuracy, empathy, information quality, and readability of HFMD-related responses generated by five publicly accessible chatbots. Methods In this exploratory cross-sectional study, 20 researcher-developed English-language prompts were constructed from authoritative public-health sources, Google Trends topic mapping, and caregiver-informed wording refinement. ChatGPT-4o, Gemini 2.5 Pro, Copilot, Doubao, and DeepSeek-V3.2-Exp were evaluated between April 2 and April 5, 2026. Five trained reviewers independently assessed responses against predefined reference standards using safety and accuracy criteria, an empathy scale, DISCERN, EQIP, the Global Quality Score, and JAMA benchmarks. Six formula-based readability indices were calculated. Paired comparisons used Friedman and Cochran Q tests, with prespecified post-hoc procedures and Benjamini-Hochberg correction. Results Unsafe-response rates ranged from 5.0 to 15.0%, with no statistically significant difference detected among chatbots (p = 0.797). No statistically significant inter-model differences were detected for accuracy, empathy, DISCERN, EQIP, JAMA, or Global Quality Score in the 20-prompt set; because the study was exploratory and was not powered to establish equivalence, these findings do not demonstrate comparable or interchangeable performance. All six readability indices differed significantly among chatbots (p = 0.030 to <0.001). ChatGPT and Doubao generally produced lower estimated grade-level complexity than Gemini and DeepSeek. Ten potentially unsafe or misleading responses were identified, mainly involving overgeneralization of EV71 vaccine protection, hand-hygiene qualification, and disinfection advice. Conclusion Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained. These systems may support general information seeking, but their responses require cautious interpretation and should not replace individualized professional advice.

Zhao-Le Gong, Yan Na, Yi Guo et al. · 0 citations
Open access 2026

Can ChatGPT be a dependable resource for parents looking for information on specific learning disorder?

ChatGPT is an innovative artificial intelligence (AI) tool that is increasingly applied in healthcare and is widely used by professionals and the public for medical information. This study evaluated ChatGPT’s capacity to provide reliable, evidence-based information on specific learning disorders (SLDs) to parents. A set of questions about specific learning disorders, which are based on questions frequently asked by patients and their families in psychiatric interviews and supported by relevant literature, was prepared by two experienced child and adolescent psychiatry specialist. Two mental health professionals independently evaluated the responses. The findings indicate that ChatGPT generally provides accurate and comprehensive information on SLDs. However, parents’ responses often included technical terminology and excessive detail, reducing clarity. Although ChatGPT has the potential to deliver accurate and detailed information, it may cause misunderstandings or undue concerns for nonspecialists. This study evaluates the accuracy, clarity, and inclusivity of ChatGPT’s responses to questions posed by different user profiles—a parent/caregiver and an experienced expert. Results indicated that ChatGPT demonstrated consistently high accuracy, with median accuracy scores of 5 and mean scores of 4.8 and 4.7 for responses given to parents/caregivers and experts, respectively. No significant difference was found between the two groups. Similarly, inclusivity scores were high across both groups, with median values of 5 and mean scores of 4.82 and 4.87, respectively, showing no significant difference. However, clarity and comprehensibility varied by user group. These findings emphasize the need for refinements to enhance ChatGPT’s adaptability to different levels of user expertise. Ensuring that responses are tailored appropriately for lay audiences can improve clarity while maintaining accuracy. Future developments should focus on optimizing language simplification and contextual adaptation, reinforcing ChatGPT’s role as a supplementary tool rather than an independent medical information source for non-expert users.

Unknown authors · 0 citations
Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

SUMMARY OBJECTIVE: Older adults increasingly use artificial intelligence-based tools to obtain health information. Although artificial intelligence chatbots such as ChatGPT may enhance access, the quality, readability, and patient safety of fall-prevention information remain uncertain. This study aimed to evaluate the quality, readability, and patient safety implications of ChatGPT-generated responses to common questions about fall risk and home safety in older adults. METHODS: Ten frequently asked fall-related questions were submitted to ChatGPT (version 5.2). Responses were independently assessed by a multidisciplinary panel including physiotherapists, a geriatrician, a physical medicine and rehabilitation physician, an occupational therapist, and an orthopedic specialist. Quality was evaluated using the Mika classification. Readability was measured with the Flesch-Kincaid Grade Level. Interrater reliability was analyzed using a two-way random-effects intraclass correlation coefficient model with absolute agreement (intraclass correlation coefficient [2,k]). RESULTS: Three responses were rated as "excellent," while seven responses were rated as "satisfactory requiring minimal clarification." No response received a rating corresponding to "moderately satisfactory" or "unsatisfactory." The mean Flesch-Kincaid Grade Level was 8.4 (range 4.3–11.9). Five responses exceeded the readability levels commonly recommended for patient education materials. Interrater reliability demonstrated fair agreement (intraclass correlation coefficient [2,k]=0.72; 95%CI 0.64–0.80). CONCLUSION: While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns. AI-generated health content should be reviewed and tailored to older adults’ health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations
Review Jul 2026

How Correct is AI for Infant Safe Sleep Advice? Evaluating Accuracy of ChatGPT, Gemini, and Claude Against AAP Guidelines.

OBJECTIVE Since the 1994 "Back to Sleep" campaign, pediatricians have promoted evidence-based infant safe sleep practices to reduce sleep-related infant deaths. However, caregivers increasingly seek guidance online. We sought to determine the accuracy of large language model (LLM) responses to caregiver questions about infant safe sleep, compared with the American Academy of Pediatrics' (AAP) 2022 recommendations. DESIGN Nine caregiver questions adapted from Reddit New Parents forum were mapped to core AAP safe sleep topics. Each was entered into three LLMs: ChatGPT 5, Gemini 2.5 Flash, and Claude Sonnet 4.5, three times within the same day to assess stability. Three reviewers scored responses on a 0-2 scale for accuracy (primary outcome), completeness, and empathy. Stability reflected similarity across repeated responses. Readability was calculated using the Flesch-Kincaid grade level. Mean scores were compared using descriptive statistics, analysis of variance, and post hoc testing. RESULTS Mean accuracy varied significantly across models. Gemini had the highest accuracy score (mean 1.85), followed by Claude (1.44), and ChatGPT (1.30). Gemini was significantly more accurate than ChatGPT (p=0.01). All models scored high in empathy (2). There were no significant differences in completeness and stability between models. ChatGPT had the lowest average readability, with all models' reading levels between grades seven to nine (7.64 vs 9.15 vs 8.82, p=0.01). Direct guideline questions yielded higher accuracy than nuanced questions. CONCLUSION LLMs offer inconsistently accurate but empathetic infant safe sleep advice with frequent deviations from AAP recommendations. Pediatric oversight and collaboration with technology developers are essential to ensure safe, evidence-based information for families.

Evin Rothschild, Jack Christian, C. Arar et al. · 0 citations
Open access Jul 2026

Evaluation of the accuracy and consistency of DeepSeek and ChatGPT in addressing type 1 diabetes mellitus-related queries in adults.

Adult type 1 diabetes mellitus (T1DM) involves complex diagnosis, treatment, and long-term self-management, creating a need for accurate and accessible health education. Large language models (LLMs) are increasingly used for medical information seeking, yet their accuracy and consistency in adult T1DM-related queries remain insufficiently evaluated. A guideline-based comparative evaluation assessed DeepSeek-V3.2 and ChatGPT-5.0 using 22 English-language prompts derived from the 2021 ADA/EASD consensus report, covering basic knowledge, diagnosis and differential diagnosis, treatment, and complications. The prompts were submitted to both models twice, two weeks apart. Responses were independently evaluated by two blinded endocrinology specialists using a predefined four-point scoring rubric, with disagreements adjudicated by a third senior endocrinologist. Short-term consistency was assessed by expert judgment and TF-IDF cosine similarity. Inter-rater agreement was good (Cohen's κ = 0.71). Expert-judged consistency was 95.45% (21/22) for both models; TF-IDF cosine similarity was 0.52 ± 0.09 for DeepSeek and 0.54 ± 0.09 for ChatGPT. Overall accuracy scores were 3.59 ± 0.59 and 3.77 ± 0.43, respectively, with no statistically significant difference (p = 0.102). Comprehensive ratings accounted for 63.64% and 77.27%, respectively, and mixed correct and incorrect or outdated information accounted for 4.55% and 0.00%. Both models may support adult T1DM-related health education, but outputs should be interpreted as supplementary educational material under professional guidance rather than as diagnostic or therapeutic advice.

Junyu Feng, Xiao Fei, Tingyi Qian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.