Skip to content

Author

D. Buonsenso

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

LLM-as-a-judge for infection prevention and control and antimicrobial resistance impact: comparing three main LLMs vs. human experts' assessment

Background Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks. Methods We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence between automated and human assessments, adhering to CHART reporting guidelines. Results Analysis of 404 evaluations revealed a systematic upward divergence: all LLMs consistently assigned higher scores than human experts. This optimism bias persisted after adjusting for domain-specific differences and clustering effects. The gap was particularly pronounced in domains of persuasiveness and AMR impact, while information quality showed more heterogeneous results. Intra-rater reliability assessments demonstrated that LLMs maintained stable scoring patterns under identical prompting conditions. Conclusions LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators. These results do not support the use of LLMs for autonomous evaluation in high-stakes public health contexts. Rather, LLM-based judging is best suited as a scalable screening tool within supervised human-in-the-loop workflows, where expert oversight serves as a necessary safeguard for evidence-based accuracy.

Marcello di Pumpo, Leonardo Villani, M. R. Gualano et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.