Large language models for breast cancer treatment planning: a blinded real-world evaluation of DeepSeek, ChatGPT, and oncologist recommendations
Abstract
Rationale and objectives Large language model (LLM) are increasingly explored for oncology decision support, yet their alignment with real-world clinical practice across varying disease complexities remains insufficiently characterized. This study aimed to evaluate and compare the accuracy, stability, and concordance of two advanced LLMs—DeepSeek V3.1 and ChatGPT-5—against experienced oncologists in generating breast cancer treatment plans within a specific clinical setting. Materials and methods This retrospective study compared the performance of DeepSeek V3.1 and ChatGPT-5 with senior oncologists using de-identified records from 213 breast cancer patients (Stages I–IV). To assess effectiveness, we implemented a multidimensional evaluation framework: accuracy was measured using a 5-point Likert scale by three independent, blinded expert reviewers; internal consistency was quantified via variance and coefficient of variation; and clinical concordance was evaluated using a structured five-level scoring system. Statistical analyses, including ANOVA and ordinal regression, were used to examine the impact of disease stage on AI-human agreement. Results Under standardized retrospective review conditions, LLM-generated recommendations demonstrated higher expert-rated guideline concordance and lower variability than historical real-world oncologist plans. Specifically, DeepSeek V3.1 achieved the highest expert-rated accuracy scores with minimal internal variance (4.91 ± 0.36), outperforming both ChatGPT-5 (4.65 ± 0.62) and clinicians (3.82 ± 0.63, P < 0.001). While AI outputs exhibited high mutual consistency (74.2%), expert evaluations revealed a significant decline in AI-clinician agreement as disease stage advanced (P < 0.001), particularly in Stage IV cases where clinicians prioritized real-world constraints such as financial toxicity. Conclusions Advanced LLMs, particularly DeepSeek V3.1, demonstrated strong performance in generating standardized, guideline-concordant breast cancer treatment plans, showing superior consistency over human specialists in protocol-driven scenarios. However, the widening gap in complex late-stage cases highlights limitations in accounting for clinical context and socioeconomic factors. These findings support the role of LLMs as robust clinician-supervised decision-support tools while emphasizing the necessity of human judgment for individualized care.