PURPOSE
To evaluate whether deep learning-based segmentation can improve visualization and diagnostic interpretation of MR cholangiopancreatography (MRCP) maximum intensity projections (MIP) by suppressing overlapping high-intensity anatomy while preserving the pancreatobiliary system.
METHODS
A total of 322 3D MRCP datasets from 162 patients were included. The training set comprised 265 cases from 127 patients, allowing multiple acquisitions per patient, whereas the evaluation set consisted of 35 cases with a single acquisition per patient. A deep learning-based segmentation model was trained using manual annotations with three distinct labels: background, primary structures, and secondary structures. Conservative safety margins were applied to reduce segmentation omissions. Two board-certified radiologists independently rated processed and original MIP images using 4-point Likert scales (1 = poor, 4 = excellent). Two-sided Wilcoxon signed-rank tests with Benjamini-Hochberg correction were used, and inter-reader agreement was assessed using linearly weighted Cohen's κ .
RESULTS
Segmentation suppressed obscuring structures and improved biliary visualization. After correction for multiple comparisons, processed images showed significantly improved structure overlap scores compared with original images (3.5 ± 0.9 vs. 2.5 ± 0.6) and higher diagnostic confidence (3.4 ± 0.7 vs. 3.2 ± 0.7).
CONCLUSIONS
Deep learning-based segmentation improves MRCP MIP visualization by reducing anatomical overlap while preserving clinically relevant pancreatobiliary anatomy.
Jinho Kim, S. Werner, S. Afat et al.· MAGMA· 0 citations
Gastrointestinal stromal tumors (GISTs) are molecularly heterogeneous neoplasms whose management depends on individualized, multidisciplinary decision-making. While multidisciplinary tumor boards (MTBs) represent the standard of care, access remains limited in many clinical settings. This study evaluates the performance of two large language models in generating GIST MTB recommendations and assesses their agreement with expert MTB decisions using predefined clinical evaluation criteria. This retrospective single-center study included 99 GIST cases discussed at an institutional MTB. A structured prompt was developed to extract clinical variables and generate treatment recommendations. ChatGPT-5 and Qwen3 were independently evaluated across five predefined domains: diagnostic recommendations, therapeutic modalities, treatment sequence and timing, systemic therapy regimen selection, and clinical contextualization. Two expert reviewers scored all outputs in a blinded fashion. Normalized scores, inter-model comparisons, perfect-case rates, and inter-rater agreement were analyzed. Both models demonstrated high concordance with expert MTB recommendations, with mean total normalized scores of 0.901 for ChatGPT-5 and 0.875 for Qwen3, without a significant difference between models (p > 0.05). Perfect agreement was observed in 52.5% of ChatGPT-5 cases and 48.5% of Qwen3 cases (p > 0.05). Diagnostic recommendations scored significantly lower than all other domains in both models (all adjusted p < 0.05). Overall inter-rater agreement was almost perfect (weighted Cohen’s kappa=0.978). Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them. Diagnostic reasoning represented the weakest domain, reflecting the challenge of reconstructing context-dependent workup decisions from tumor board documentation. These findings support a potential assistive role for LLMs in GIST MTB workflows, while underscoring the continued necessity of expert oversight.