When Artificial Intelligence Speaks For The Obstetrician: Multilingual Accuracy On Real Patient Questions
BackgroundThe use of generative large language models (LLMs) in healthcare is rapidly increasing, offering easier access to medical information. However, comprehensive data on their multilingual accuracy and the reliability of cited scientific references remain limited. This study aimed to compare the response accuracy and reference quality of three free LLMs (ChatGPT, Google Gemini, DeepSeek) in Turkish and English using common pregnancy-related questions.MethodsIn this comparative observational study, 14 frequently asked pregnancy questions were posed to each LLM in Turkish and English, requesting responses supported by up-to-date scientific web sources. Answers were evaluated blindly by obstetricians and gynaecologists for accuracy. References were independently assessed for reliability, scientific validity, and accessibility. Statistical analyses were performed.ResultsLanguage and model infrastructure significantly influenced performance. Google Gemini and DeepSeek provided more accurate responses in English than in Turkish (p