Skip to content

Assessing the performance of three large language models in thyroid cancer tumour board decision-making.

Aug 2026 · Updates in Surgery · 0 citations · 13 references
Medicine

TL;DR

ChatGPT achieved the highest overall concordance, although all models generated clinically acceptable recommendations in most cases, and all LLMs demonstrated high concordance with consultant-led thyroid cancer MDT decisions.

View source

Similar papers

#large language models Open access Sep 2026

Large Language Models Versus Multidisciplinary Tumor Board Decisions in Thyroid Cancer

ABSTRACT Objectives Large language models (LLMs) are increasingly proposed as clinical decision‐support tools; however, their agreement with real‐world multidisciplinary tumor board (MDT) decisions remains insufficiently investigated in thyroid oncology. To evaluate the concordance between treatment recommendations generated by ChatGPT 5.2 and Gemini 3.0 and decisions made by a tertiary multidisciplinary thyroid tumor board. Methods This study included 59 consecutive patients discussed at a tertiary MDT between August and December 2025. Anonymized clinical data, including demographics, ultrasonographic findings, and Bethesda cytology, were provided to both LLMs using standardized structured prompts. MDT decisions were defined as the reference standard. Agreement was assessed using exact concordance rates and Cohen's kappa (κ) statistics with 95% confidence intervals. Results ChatGPT 5.2 achieved a concordance rate of 71.2% (42/59), demonstrating substantial agreement (κ = 0.623; 95% CI 0.459–0.771). Gemini 3.0 showed a concordance rate of 64.4% (38/59), reflecting moderate agreement (κ = 0.527; 95% CI 0.359–0.684). Discordance increased in complex scenarios involving lateral neck dissection, radioactive iodine therapy, and active surveillance. Conclusions While LLMs demonstrate promising concordance in standardized thyroid cancer management, they are best positioned as supportive decision aids—such as in MDT preparation and workflow streamlining—rather than replacements for expert multidisciplinary evaluation, particularly in complex clinical scenarios. Level of Evidence 3.

Bayram Barış Büyük, Arzu Or Koca, Felat Toprak et al. · 0 citations
Review Open access Aug 2026

Precision oncology meets Generative AI: assessing large language models in multidisciplinary GIST tumor boards

Gastrointestinal stromal tumors (GISTs) are molecularly heterogeneous neoplasms whose management depends on individualized, multidisciplinary decision-making. While multidisciplinary tumor boards (MTBs) represent the standard of care, access remains limited in many clinical settings. This study evaluates the performance of two large language models in generating GIST MTB recommendations and assesses their agreement with expert MTB decisions using predefined clinical evaluation criteria. This retrospective single-center study included 99 GIST cases discussed at an institutional MTB. A structured prompt was developed to extract clinical variables and generate treatment recommendations. ChatGPT-5 and Qwen3 were independently evaluated across five predefined domains: diagnostic recommendations, therapeutic modalities, treatment sequence and timing, systemic therapy regimen selection, and clinical contextualization. Two expert reviewers scored all outputs in a blinded fashion. Normalized scores, inter-model comparisons, perfect-case rates, and inter-rater agreement were analyzed. Both models demonstrated high concordance with expert MTB recommendations, with mean total normalized scores of 0.901 for ChatGPT-5 and 0.875 for Qwen3, without a significant difference between models (p > 0.05). Perfect agreement was observed in 52.5% of ChatGPT-5 cases and 48.5% of Qwen3 cases (p > 0.05). Diagnostic recommendations scored significantly lower than all other domains in both models (all adjusted p < 0.05). Overall inter-rater agreement was almost perfect (weighted Cohen’s kappa=0.978). Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them. Diagnostic reasoning represented the weakest domain, reflecting the challenge of reconstructing context-dependent workup decisions from tumor board documentation. These findings support a potential assistive role for LLMs in GIST MTB workflows, while underscoring the continued necessity of expert oversight.

Reza Dehdab, Judith Herrmann, Fiona Mankertz et al. · 0 citations
Open access Jul 2026

Performance of leading large language models in adhering to clinical guidelines for anaplastic thyroid cancer: a comparative study

Leading LLMs show variable capacity to align with ATC clinical guidelines, while top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use.

Mohamed Yasser, Ghada Barakat, S. Awny et al. · 0 citations
Open access Jul 2026

A multidimensional benchmarking framework for large language models in oncologic decision making

A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.

M. Halıcı, Serkan Saltürk, Irem Sayin et al. · 0 citations
Aug 2026

Comparative Performance of Multimodal Large Language Models in Grayscale Ultrasound-Based Classification of Thyroid Nodules.

BACKGROUND Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear. OBJECTIVE This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules. METHODS This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison. RESULTS All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs. CONCLUSION Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.

Ziman Chen, Yingli Wang, Fei Chen · 0 citations
Review Jul 2026

Correct but Incomplete: Limitations of AI-Assisted Decision Support in Rectal Cancer.

BackgroundArtificial intelligence (AI) increasingly supports clinical decision making. ChatGPT-5 (OpenAI, San Francisco, CA) and OpenEvidence (OpenEvidence Inc, Miami, FL) are regularly used by physicians, yet the limits of their reliability in decision-making remain poorly defined. This study evaluated both platforms against the National Comprehensive Cancer Network® (NCCN) Guidelines for Rectal Cancer to define where AI tools perform well and where they fall short.MethodsThe NCCN Guidelines for Rectal Cancer (Version 2.2025) were reviewed. Three questions were generated per decision-making page and classified into workup/diagnosis, treatment, and surveillance domains, yielding 138 clinical scenarios. Both platforms were queried. Responses were scored independently by two physicians on a 5-point Likert scale (5 = completely correct; 1 = absolutely incorrect). Two performance thresholds were defined: Correctness (≥3) and Accuracy (≥4). Proportions were compared with Fisher's exact test and score distributions with the Mann-Whitney U test.ResultsBoth platforms demonstrated high overall guideline concordance. ChatGPT-5 achieved Correctness in 136 (98.6%) and Accuracy in 116 (84.1%), compared to 128 (92.8%) and 112 (81.2%) for OpenEvidence, respectively. Overall Correctness favored ChatGPT-5 (P = 0.035), while Accuracy showed no significant difference (P = 0.634). Mean Likert scores were 4.57 vs 4.37 (P = 0.089). Both platforms achieved >90% Correctness across all domains.ConclusionBoth platforms demonstrate reliable identification of the broad direction of care but exhibit important limitations in completeness, nuance, and the handling of preference-sensitive decisions. These findings define the appropriate role of AI-assisted decision support as an adjunct to, rather than a substitute for, multidisciplinary review in rectal cancer management.

Ryan J Meyer, Tamir E. Bresler, Tadevos T Makaryan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.