1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jun 2026

Large language models for breast cancer treatment planning: a blinded real-world evaluation of DeepSeek, ChatGPT, and oncologist recommendations

Rationale and objectives Large language model (LLM) are increasingly explored for oncology decision support, yet their alignment with real-world clinical practice across varying disease complexities remains insufficiently characterized. This study aimed to evaluate and compare the accuracy, stability, and concordance of two advanced LLMs—DeepSeek V3.1 and ChatGPT-5—against experienced oncologists in generating breast cancer treatment plans within a specific clinical setting. Materials and methods This retrospective study compared the performance of DeepSeek V3.1 and ChatGPT-5 with senior oncologists using de-identified records from 213 breast cancer patients (Stages I–IV). To assess effectiveness, we implemented a multidimensional evaluation framework: accuracy was measured using a 5-point Likert scale by three independent, blinded expert reviewers; internal consistency was quantified via variance and coefficient of variation; and clinical concordance was evaluated using a structured five-level scoring system. Statistical analyses, including ANOVA and ordinal regression, were used to examine the impact of disease stage on AI-human agreement. Results Under standardized retrospective review conditions, LLM-generated recommendations demonstrated higher expert-rated guideline concordance and lower variability than historical real-world oncologist plans. Specifically, DeepSeek V3.1 achieved the highest expert-rated accuracy scores with minimal internal variance (4.91 ± 0.36), outperforming both ChatGPT-5 (4.65 ± 0.62) and clinicians (3.82 ± 0.63, P < 0.001). While AI outputs exhibited high mutual consistency (74.2%), expert evaluations revealed a significant decline in AI-clinician agreement as disease stage advanced (P < 0.001), particularly in Stage IV cases where clinicians prioritized real-world constraints such as financial toxicity. Conclusions Advanced LLMs, particularly DeepSeek V3.1, demonstrated strong performance in generating standardized, guideline-concordant breast cancer treatment plans, showing superior consistency over human specialists in protocol-driven scenarios. However, the widening gap in complex late-stage cases highlights limitations in accounting for clinical context and socioeconomic factors. These findings support the role of LLMs as robust clinician-supervised decision-support tools while emphasizing the necessity of human judgment for individualized care.

Ming Li, Yiran Yu, Gang Li et al. · 0 citations