Comparative performance of ChatGPT-4.0, DeepSeek, Gemini and Perplexity in answering common questions from patients with COPD
Abstract Large language models (LLMs) may help with patient education, but clinicians must check LLM responses for accuracy, comprehensiveness, readability, and safety. This study compared responses from the LLMs ChatGPT-4.0, DeepSeek, Gemini, and Perplexity to 20 English-language questions about COPD that patients could present. Each question was submitted once to each LLM. Three COPD experts were blinded to the source LLM and independently scored 80 LLM responses for accuracy and comprehensiveness. Three readability indices were recorded; because they were strongly correlated, Flesch-Kincaid Grade Level (FKGL) was used as the primary readability metric. Accuracy differed across responses from the four LLMs (Friedman χ²(3) = 38.902, P < 0.001, Kendall’s W = 0.648), as did comprehensiveness (χ²(3) = 52.468, P < 0.001, Kendall’s W = 0.874). ChatGPT-4.0 and DeepSeek did not differ significantly on either outcome, and both scored higher than Gemini and Perplexity. Mean ratings from three experts showed moderate inter-rater reliability for accuracy [ICC(A,3) = 0.571] and comprehensiveness [ICC(A,3) = 0.685]. Perplexity had the highest FKGL, and no LLM consistently met the approximate eighth-grade target. Patient-facing LLM responses require clinical review.