Skip to content
Open access

Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health.

Jul 2026 · npj Digital Medicine · 1 citation
Medicine

TL;DR

Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.

Abstract

Evaluating generative AI output remains a critical bottleneck for safe and scalable deployment of AI in healthcare. Expert clinical judgement is often presented as the gold standard, but human assessment is costly and inconsistent. LLM-as-judge systems, i.e., leveraging AI to evaluate other AI outputs, have been proposed, yet their reliability in global health remains untested. We compared five LLM judges and six human clinicians in evaluating responses to questions posed by Rwandan health workers. The highest-performing LLM-judge (Claude-4.1-Opus) matched human evaluators on only four of eleven evaluation criteria, with other models scoring too leniently (Gemini-2.5-Pro) or too harshly (GPT-5). Constructing LLM-juries to balance model-specific biases improved agreement on only one additional criterion. Notably, performance and cost-effectiveness fell when moving from English to Kinyarwanda. Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.

Read PDF

Similar papers

Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care.

Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations
Jul 2026

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

Readiness stress-testing of medical AI is extended to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement.

Koyar Afrasyab · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes: (1) base, using the instruct model as is; (2) SFT, distillation-based supervised fine-tuning only; (3) RL, GRPO-based reinforcement learning only; and (4) SFT$\rightarrow$RL, SFT followed by RL. The base and single-stage regimes struggle on structured medical discrimination such as PICO extraction and clinical calculations, whereas SFT$\rightarrow$RL performs best across correctness, compliance, and robustness; gains concentrate on decomposable tasks (PICO, MedCalc), at times matching or outperforming frontier models.

R. de Oliveira, Federico Pittino, J. Gwinnutt et al. · 0 citations
Review Open access Aug 2026

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting of patient-initiated text messages in a multi-state Medicaid population.

Sanjay Basu, Sadiq Y. Patel, Parth Sheth et al. · 0 citations
Open access Jul 2026

Comparing Human and AI-Assisted Evaluation of Interpreting Performance

Growing data volumes and the rise of AI-assisted interpreting solutions increasingly challenge the assessment of interpreting quality. This study investigates whether large language models (LLMs) can support error-based interpreting performance assessment, which has traditionally relied on human-driven approaches. We used transcripts from 42 English-to-Polish interpreting outputs (Broś et al. 2025), evaluated using the NTR model (Romero Fresco & Pöchhacker 2017). Two trained human evaluators conducted detailed analyses, which were then compared against ChatGPT-4o outputs generated through different prompting strategies. Initial results showed an overlap of up to 85% in omission detection between human and AI assessments when a full prompting sequence was applied and the source and target texts were provided in the chat's interface rather than in uploaded text files. However, bulk and iterative analyses revealed significant limitations, including prompt drift and reduced accuracy over successive tasks, with the lowest overlaps below 10%. Despite these drawbacks, the AI model occasionally identified errors overlooked by humans and demonstrated potential for expediting error annotation. The study highlights that, while LLMs provide partial automation and enhance consistency, comprehensive human oversight remains indispensable. Ultimately, integrating AI with human expertise can become a promising hybrid approach to interpreting quality evaluation, balancing efficiency with nuanced judgment.

Tomasz Korybski, Karolina Broś, Małgorzata Szupica-Pyrzanowska · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.