Skip to content
Review

HealMed: Multilingual Evaluation of Large Language Models in Medicine

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models, whereas many open-source and medically specialized models showed larger and less consistent gaps.

Abstract

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.

View source

Similar papers

Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions

Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.

Minh Duc Bui, Mario Sanz-Guerrero, Abteen Ebrahimi et al. · 0 citations
Open access Aug 2026

When Artificial Intelligence Speaks For The Obstetrician: Multilingual Accuracy On Real Patient Questions

BackgroundThe use of generative large language models (LLMs) in healthcare is rapidly increasing, offering easier access to medical information. However, comprehensive data on their multilingual accuracy and the reliability of cited scientific references remain limited. This study aimed to compare the response accuracy and reference quality of three free LLMs (ChatGPT, Google Gemini, DeepSeek) in Turkish and English using common pregnancy-related questions.MethodsIn this comparative observational study, 14 frequently asked pregnancy questions were posed to each LLM in Turkish and English, requesting responses supported by up-to-date scientific web sources. Answers were evaluated blindly by obstetricians and gynaecologists for accuracy. References were independently assessed for reliability, scientific validity, and accessibility. Statistical analyses were performed.ResultsLanguage and model infrastructure significantly influenced performance. Google Gemini and DeepSeek provided more accurate responses in English than in Turkish (p

A. Tığlı, Y. Baykuş, R. Deniz et al. · 0 citations
Preprint Aug 2026

PlainMedScale: A Corpus of Multi-Level Simplified Medical Texts in German and English

We introduce PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional and consumer), Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS. The four tiers correspond to distinct communicative functions --- reference, explanation, decision support, and access --- and move beyond the binary expert--lay contrast of prior corpora. In two pilot studies enabled by the alignments, we show that many readability metrics established on two registers fail to generalize across the full gradient, and that a SOTA open-weight LLM prompted for Plain Language still partially preserves the difficulty of its input. Code (https://github.com/GS-Uni-Heidelberg/PlainMedScale) and data (https://doi.org/10.5281/zenodo.21728290) are made available.

Bruno Brocai, Ilaria Papagno, Mayumi Ohta · 0 citations
Review Open access Aug 2026

Performance of Large Language Models in Answering Otolaryngology Clinical Questions in Non‐English Languages

ABSTRACT Background Large language models (LLMs) such as ChatGPT are being explored for various medical applications, but their performance across languages is not established. Despite the promising results in English from previous studies, it is unclear how ChatGPT performs in other languages. We assessed the ability of ChatGPT‐4 to answer common otolaryngology‐related frequently asked questions in Chinese, Korean, and Spanish in this study. Methods Eighteen clinical questions were selected from the American Academy of Otolaryngology‐Head and Neck Surgery Foundation guidelines and ENT Health across six subspecialties. Bilingual clinicians translated the questions to Chinese, Korean, or Spanish and input them into ChatGPT‐4 (OpenAI, San Francisco, USA). The responses were rated on a five‐point Likert scale for accuracy, completeness, and similarity to English reference answers. A linear mixed‐effects model with language as a fixed effect and random intercepts for question and rater was used to compare performance across languages. Inter‐rater reliability was assessed using intraclass correlation coefficients and Krippendorff's α. Subjective evaluations were also collected. Results ChatGPT‐4's performance differed significantly by language in all assessment domains. Responses in Spanish scored highest throughout and most frequently were characterized as clear and organized. Korean answers were generally accurate but tended to be more conservative and less detailed. Chinese answers showed greater variability, with reviewers noting a mix of clear and less accurate outputs. Conclusion The clinical performance of ChatGPT‐4 differed across Chinese, Korean, and Spanish. These findings underscore the importance of ongoing evaluation and refinement of LLMs to ensure they provide reliable multilingual support in clinical practice.

Eugene Oh, Arthur W. Wu, Laura García-Rodríguez et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.