Feasibility and expert evaluation of a medical large language model in an interdisciplinary clinical board
Abstract
Interdisciplinary clinical boards (ICB) on chronic inflammatory diseases discuss complicated and multidisciplinary cases with inflammatory conditions among experts in their discipline. Medically specialized large language models (LLMs) are finding increasingly widespread use among physicians to answer clinical questions. LLMs offer the potential to close the gap between rapidly expanding medical knowledge and limited time of physicians by providing evidence-based summaries to clinical questions and enable “real-time” literature search. However, it is unclear if LLMs can be implemented in clinical routine of ICB. In this prospective study, 100 consecutive real-world cases were discussed in our ICB. Each case was first discussed independently of the LLM and a preliminary recommendation was documented. The clinical question was then entered in the LLM “OpenEvidence”, the answer was reviewed by the board with the opportunity to add aspects not previously raised, and the recommendation was subsequently confirmed as final. Each participating expert evaluated the answer of the LLM using a questionnaire comprising eight Likert scale (1–5) items and one free-text item. Two physicians, blinded to each other, additionally evaluated outputs for their agreement with the board’s decision. The same cases were evaluated after prompt optimization at two different time points. The 100 discussed cases resulted in 759 rater-case evaluations. A mean of 2.3 specialties were directly involved in each case. The highest score was achieved by the item “OpenEvidence provided clinically relevant information for this case.” (4.06). Recommendation for future use achieved a mean score of 3.95. Complete agreement of the LLM answer with the final recommendation of the ICB was rated in 40.2% and 50.5% of all evaluable cases (n = 97) by the two raters. When restricted to consensus cases, i.e. cases rated identically by both raters (n = 62 of 97), complete agreement was 51.6%. Outputs after prompt optimization tended to be associated with higher agreement, but did not reach statistical significance. Outputs were not stable over time. The integration of OpenEvidence into an ICB was feasible and was perceived as useful by the participating specialists, with partial or complete concordance with the final board recommendation in the majority of evaluable cases. As the study did not assess patient outcomes, clinical harm or hallucinations, no conclusions on patient safety can be drawn. The LLM was naturally limited by the information provided in the prompt. Inadequate questions or missing information led to suboptimal answers, highlighting the importance of identifying and including relevant information beforehand. Expert oversight remains essential. Model drift of LLMs may outweigh prompt optimization.