Evaluating Large Language Models on Clinical Risk Judgement: An Example from Alcohol Use Disorder
Large language models (LLMs) are being explored for medical risk assessment, yet performance on standardized benchmarks does not necessarily establish clinical competency. The present study evaluated whether LLM-generated Alcohol Use Disorder (AUD) risk judgments aligned with epidemiological evidence and remained consi...