Author

Shih-Chi Ku

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care

Large language model (LLM)-as-a-judge systems offer scalable evaluation of artificial intelligence (AI)-generated clinical outputs, yet their susceptibility to prompt variability raises concerns regarding reproducibility and alignment with expert judgement. This study examined whether evaluation prompt strategies influence scoring patterns and concordance with clinical raters in critical care. This post-hoc analysis used 90 structured clinical reports generated in a prior study using an XGBoost ICU mortality prediction model trained on the MIMIC-IV database. GPT-4o (Azure AI, version 2024-11-20) produced structured interpretations from risk estimates and SHAP attributions. These outputs were evaluated using the IMPACT framework under three evaluation prompt strategies: baseline (E1), top-down decremental (E2), and bottom-up incremental (E3). Agreement between clinician ratings and the automated o3-mini evaluator (Azure AI, version 2025-01-31) was assessed using intraclass correlation coefficients (ICC), with strategy comparisons by Fisher’s z-transformation. Score deviations were examined with repeated-measures ANOVA. Mean IMPACT scores were 79.9 (SD 9.9) for E1, 83.3 (SD 9.6) for E2, 78.7 (SD 9.1) for E3, and 78.6 (SD 8.9) for clinicians. All strategies demonstrated substantial agreement (ICC > 0.80). E2 showed significantly lower agreement with clinicians (ICC = 0.82) than E1 and E3 (both ICC = 0.94, p < 0.001). Score deviations differed significantly across strategies (p < 0.001), with E3 showing the smallest mean deviation (0.1) and E2 the largest (4.7). Prompt design meaningfully affects both IMPACT scoring patterns and the reliability of LLM-based evaluators. Bottom-up incremental scoring showed the closest alignment with human assessment, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations