Open access
Jul 2026
Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health.
Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.
G. Williams, S. Rutunda, Floris Nzabakira et al.
· npj Digital Medicine · 1 citation