Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias.
Abstract
Language barriers hinder healthcare, particularly during case history-taking, a key part of diagnosis. While multilingual artificial intelligence (AI) chatbots offer solutions, there is fragmented evidence of their effectiveness and impact. This systematic review followed PRISMA 2020 guidelines, examining studies published between 2015 and 2025 on multilingual AI chatbots in healthcare across four databases (Google Scholar, Scopus, Web of Science, and PubMed), using a two-stage screening process. Data extraction focused on applications, supported languages, underlying technologies, target populations, and clinical outcomes. From 503 records, 49 studies, covering primary care, telemedicine, oncology, mental health, and other areas, met the criteria. Supported languages included English, Spanish, Arabic, Chinese, Hindi, and other underrepresented languages. In individual system evaluations using heterogeneous methodologies and evaluation settings, AI chatbots achieved a diagnostic accuracy ranging from 72–92%. Core technologies included large language models (LLMs), bidirectional encoder representations from transformers (BERT), a generative pre-trained transformer (GPT), retrieval-augmented generation (RAG), speech recognition, and distillation. The findings show that these improve clinical workflow (30–70% time savings) and patient engagement, reduce language barriers, and promote health equity. However, the overall evidence certainty was low to moderate, reflecting the predominance of prototype and proof-of-concept studies. Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias. Future research should include real-world studies, diverse populations, standardized outcome measures, and long-term equity assessments.
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.
Ilse Super, Olya Rezaeian, Onur Asan· International Journal of Med...· 0 citations
This study provides the first prototype of an AI-driven chatbot specifically designed for MAT professionals, demonstrating feasibility of integrating advanced AI technologies to address information access barriers in addiction treatment.
Sandra C. Nwobi, Zainab Loukil, Abbas Jawahar· Frontiers in Digital Health· 0 citations
Artificial intelligence, particularly AI-based chatbots, has gained increasing attention as a tool to support pharmacists in providing drug information services, especially in settings with pharmacist shortages. This study compared the competencies of ChatGPT, ChatGPT Pro, Perplexity, and Perplexity Pro in responding to drug-related clinical questions, and examined the effect of an Enhanced Task Translation (ETT) technique on chatbot performance, after English technical terms were embedded with Thai-language prompts. An experimental comparative design was employed. Two sets of Thai-language multiple-choice questions (MCQs), each comprising 120 items based on the Thai Pharmacy Licensing Examination, were administered: one standard set and one ETT set. Cardiology-focused clinical case scenarios were additionally presented using a SOAP-note format. Performance was assessed using a rubric adapted from the American Society of Health-System Pharmacists (ASHP) guidelines across four domains: question classification, source citation, evidence application, and communication. All four chatbots surpassed the 60% passing threshold: 72.08–83.33% for the standard set and 73.75–80.83% for the ETT set. Perplexity Pro demonstrated the highest MCQ performance, while the inclusion of English technical terms did not consistently improve results across models. In the clinical case assessment, ChatGPT Pro achieved the highest rubric score, showing strong evidence use and clinical reasoning. These findings suggest that AI chatbot competency in drug-related queries is broadly comparable to that of pharmacists. The ETT technique did not produce notable performance differences between Thai-only and mixed-language prompts. AI chatbots may support pharmacists by generating initial structured responses to drug information requests. The proposed evaluation framework provides a potential model for assessing AI competency in non-English settings and underscores the importance of jointly evaluating knowledge accuracy and clinical judgment to ensure safe integration into drug information services.
Inthira Kanchanaphibool, Panyanat Aonpong, Thanaphat Dabngoen et al.· Thai Bulletin of Pharmaceuti...· 0 citations
Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots—ChatGPT, Google Gemini, and Microsoft Copilot—when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic.
Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation–Diabetes and Ramadan (IDF-DAR) Guidelines using a 0–2 accuracy scale, a 0–4 completeness scale, and a 0–3 safety scale.
Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness (
p
= 0.068) and a language effect on safety (
p
= 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20,
p
= 0.009) and completeness (κ = 0.29,
p
= 0.001) and showed a very low κ in the safety scale (κ = 0.02,
p
= 0.81).
AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
S. Alomair, Maryam Alsuwayq, Walla Alabbad et al.· Frontiers in Medicine· 0 citations
Hospitals frequently face challenges in delivering timely and accessible information to patients due to high inquiry volumes, language barriers, and limited staff availability. This paper proposes a multilingual, voice-enabled hospital chatbot that provides real-time assistance through both text and speech interfaces. The proposed system leverages Large Language Model (LLM)-based sentence embeddings using Sentence-BERT for semantic similarity-driven question answering, along with Google Translate API for multilingual support and Google Text-toSpeech for voice responses. The chatbot supports multiple Indian languages, including English, Hindi, Telugu, Tamil, Kannada, and Marathi, enabling inclusive communication across diverse user groups. Designed as a Flask-based web application with a responsive Bootstrap interface, the system aims to achieve effective contextual understanding, reduced response time, and improved accessibility when compared to traditional rule-based hospital inquiry systems. The proposed approach highlights the potential of LLM-driven semantic retrieval-based conversational agents in enhancing patient engagement and improving operational efficiency in healthcare environments.
Kamisetty Mythri Sridevi, Kallagunta Srividhya· International Journal of Eng...· 0 citations
INTRODUCTION
Telehealth is a strategic component of primary health care and has advanced in Brazil through the National Telehealth Program. Its benefits can be enhanced by artificial intelligence (AI), which has emerged as a promising tool. This study aims to compare the performance of real human and AI-generated responses to queries submitted to the teleconsultation services of the Telehealth Center of the UFMG Faculty of Medicine (NUTEL FM-UFMG), a member of the Telehealth Brazil Program.
METHODS
This is a comparative cross-sectional study of 180 real human and AI-generated responses, evaluated in a blinded manner according to quality criteria (medical adequacy, conciseness, coherence, and comprehensibility), risk potential, authorship identification accuracy, and inquiry resolution. Data from NUTEL FM-UFMG (January 2020 to May 2024) were utilized, covering cardiology, endocrinology, and obstetrics/gynecology (OB-GYN). Statistical analysis included the Shapiro-Wilk test, Kruskal-Wallis test, Nemenyi multiple comparison test, chi-square test, and Fisher's exact test.
RESULTS
Across all specialties, a significant difference was observed in comprehensibility, with AI mean scores surpassing those of humans. For the remaining quality criteria, as well as for risk potential and inquiry resolution, no significant differences were found, despite AI scoring higher than humans. Within specific specialties, significant differences was observed in endocrinology (except conciseness) and cardiology (in conciseness); AI showed superior means. Across all specialties, as well as individually within endocrinology and OB-GYN, the accuracy of authorship identification (human vs. AI) was statistically significant.
CONCLUSION
Despite existing limitations, AI demonstrates substantial potential as a support tool for teleconsultation services.
Gabriela Dário Mendes Barros, Carlos Eduardo Menezes Amaral, César Macieira et al.· Telemedicine journal and e-h...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.