Review
Aug 2026
HealMed: Multilingual Evaluation of Large Language Models in Medicine
On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models, whereas many open-source and medically specialized models showed larger and less consistent gaps.
Yingjian Chen, Fan Gao, Sherry T. Tong et al.
· 0 citations