Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.
INTRODUCTION This study evaluated concordance between treatment-pathway outputs generated from structured clinical text by the large language model (LLM) GPT-5.4 Thinking and multidisciplinary team (MDT) decisions for mid- and low rectal cancer. MATERIALS AND METHODS This single-center retrospective study included 260 patients who underwent standardized assessment, MDT discussion, and curative-intent surgery between January 2022 and December 2023. After database lock, structured de-identified clinical text was entered into GPT-5.4 Thinking using prespecified preoperative and postoperative templates, with MDT decisions treated as real-world reference decisions. The primary endpoint was preoperative concordance; secondary endpoints included postoperative and overall concordance. Concordance metrics, directional discordance, baseline comparators, repeatability in a 50-case subset, and exploratory logistic regression models were assessed. RESULTS Preoperative, postoperative, and overall concordance rates were 57.7%, 66.9%, and 46.9%, respectively; Cohen's κ values were 0.269 and 0.449 for the preoperative and postoperative stages. Preoperatively, directional discordance relative to MDT decisions included 48 potential under-intensification, 28 potential over-intensification, and 34 directionally indeterminate or heterogeneous cases. Compared with the majority-class baseline, the LLM had the same preoperative crude concordance but higher κ and balanced category-specific concordance; postoperatively, it exceeded the majority-class baseline across these metrics. Non-identical mapped categories across three runs occurred in 12/50 preoperative and 8/50 postoperative assessments despite identical inputs. CONCLUSION GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited. Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs.