Large Language Models on EDAIC: Supplementary Material.
This repository contains the code, question bank pipeline, and statistical analysis scripts used to evaluate the accuracy and explanation quality of three large language models (Claude Sonnet 5, GPT-5.6 Terra, and Gemini 3.5 Flash) on a True/False question bank spanning eleven topic areas relevant to anaesthesiology and critical care practice (Physiology, Pharmacology, Anatomy, Physics, Statistics, General Anaesthesia, Regional Anaesthesia, Special Anaesthesia, Intensive Care, Internal Medicine, and Emergency Medicine). Each question consists of a stem paired with five independent statements (A–E); models were prompted to judge the truth value of each statement and provide a brief explanation. Model outputs were scored for correctness against a reference answer key (1,310 items total) and, for a representative subsample (495 items), independently rated by three blinded physician raters across five explanation-quality axes: scientific accuracy, completeness, reasoning quality, potential harm, and bias.