Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.
Abstract
Evaluating generative AI output remains a critical bottleneck for safe and scalable deployment of AI in healthcare. Expert clinical judgement is often presented as the gold standard, but human assessment is costly and inconsistent. LLM-as-judge systems, i.e., leveraging AI to evaluate other AI outputs, have been proposed, yet their reliability in global health remains untested. We compared five LLM judges and six human clinicians in evaluating responses to questions posed by Rwandan health workers. The highest-performing LLM-judge (Claude-4.1-Opus) matched human evaluators on only four of eleven evaluation criteria, with other models scoring too leniently (Gemini-2.5-Pro) or too harshly (GPT-5). Constructing LLM-juries to balance model-specific biases improved agreement on only one additional criterion. Notably, performance and cost-effectiveness fell when moving from English to Kinyarwanda. Overall, while LLM-judges show promise, their inability to handle linguistic and cultural context is a critical limitation, underscoring the need for further investment in scalable evaluation solutions.
Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.
Readiness stress-testing of medical AI is extended to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement.
We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes: (1) base, using the instruct model as is; (2) SFT, distillation-based supervised fine-tuning only; (3) RL, GRPO-based reinforcement learning only; and (4) SFT$\rightarrow$RL, SFT followed by RL. The base and single-stage regimes struggle on structured medical discrimination such as PICO extraction and clinical calculations, whereas SFT$\rightarrow$RL performs best across correctness, compliance, and robustness; gains concentrate on decomposable tasks (PICO, MedCalc), at times matching or outperforming frontier models.
R. de Oliveira, Federico Pittino, J. Gwinnutt et al.· 0 citations
LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates.
No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting of patient-initiated text messages in a multi-state Medicaid population.
Sanjay Basu, Sadiq Y. Patel, Parth Sheth et al.· BMC Medical Informatics and...· 0 citations
Growing data volumes and the rise of AI-assisted interpreting solutions increasingly challenge the assessment of interpreting quality. This study investigates whether large language models (LLMs) can support error-based interpreting performance assessment, which has traditionally relied on human-driven approaches. We used transcripts from 42 English-to-Polish interpreting outputs (Broś et al. 2025), evaluated using the NTR model (Romero Fresco & Pöchhacker 2017). Two trained human evaluators conducted detailed analyses, which were then compared against ChatGPT-4o outputs generated through different prompting strategies. Initial results showed an overlap of up to 85% in omission detection between human and AI assessments when a full prompting sequence was applied and the source and target texts were provided in the chat's interface rather than in uploaded text files. However, bulk and iterative analyses revealed significant limitations, including prompt drift and reduced accuracy over successive tasks, with the lowest overlaps below 10%. Despite these drawbacks, the AI model occasionally identified errors overlooked by humans and demonstrated potential for expediting error annotation. The study highlights that, while LLMs provide partial automation and enhance consistency, comprehensive human oversight remains indispensable. Ultimately, integrating AI with human expertise can become a promising hybrid approach to interpreting quality evaluation, balancing efficiency with nuanced judgment.
Tomasz Korybski, Karolina Broś, Małgorzata Szupica-Pyrzanowska· The Journal of Specialised T...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.