Strategies for Deploying Large Language Models for Ascertaining Clinical Outcomes and Sites of Metastases From Radiology Impressions in Patients With Cancer.
Jul 2026· JCO Clinical Cancer Informatics· Vol 10 3, pp.
e2500164
· 0 citations· 15 references
Medicine
TL;DR
Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text, suggesting that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.
Abstract
Purpose
To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes.
Methods
Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139.
Results
This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions).
Conclusion
Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.
The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were>0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.
J. Vences, W. Tran, N. Gimpaya et al.· 0 citations
Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.
Michael Fiorino, Mei Hainline, Tanay Poddar et al.· Neuro-Oncology Advances· 0 citations
Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Here, we conducted a multicentre cohort study using 20,277 authentic clinical spine radiological reports from three teaching hospitals in China, systematically comparing the diagnostic performance, output consistency, and generalisability of four LLMs—GPT-4o, Claude-4, Qwen-3 Max, and DeepSeek-V3.1—across nine modality–region combinations under two input modes (with-option and without-option). All models showed high overall diagnostic performance, with specificity exceeding 90% and negative predictive value exceeding 96%. Cross-institutional validation demonstrated stable recall generalisability, with recall coefficients of variation ranging from 0.7% to 5.0%. However, performance was uneven across the disease spectrum: for low-prevalence conditions, precision declined by 19–42 percentage points, indicating a persistent long-tail diagnostic deficit. Input mode and prompt formulation also produced model-specific shifts in diagnostic behaviour. These findings suggest that LLMs may support report-based spine imaging diagnosis as clinical assistive tools, but deployment should account for disease prevalence, prompt sensitivity, and the need for domain-specific optimisation.
Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al.· npj Digital Medicine· 0 citations
A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.
M. Halıcı, Serkan Saltürk, Irem Sayin et al.· Scientific Reports· 0 citations
Breast cancer remains one of the most common and life threatening cancers worldwide, and early detection is strongly associated with improved survival and reduced treatment burden. This study investigates the ability of Large Language Models to perform diagnostic prediction from structured breast cancer related data. We systematically evaluated 12 LLMs across three public datasets with different clinical characteristics: the Wisconsin Breast Cancer Dataset (WBCD) based on cytological features, the Breast Cancer Coimbra Dataset (BCCD) based on metabolic biomarkers, and the Mammographic Mass Dataset (MMD) based on mammographic attributes. The evaluation covered 13 prompting strategies, including three zero-shot variants, few-shot prompting, and multiple Chain-of-Thought (CoT) and knowledge-enhanced reasoning settings. Performance was assessed using confusion-matrix-based metrics, including accuracy, precision, recall, F1-score, specificity, and Matthews Correlation Coefficient. The results showed that performance was strongly dependent on both dataset type and prompting design, and no single model dominated all tasks. The best model–strategy pairvaried by dataset: Cogito-v1-preview-qwen-32B achieved the highest F1-score on WBCD with an F1-score of 92.00%, GPT 4.1 and GPT 4o on BCCD with an F1-score of 85.39%, and Gemini 2.5 Flash Lite on MMD with an F1-score of 82.91%. Prompt engineering had a substantial effect on outcomes, but its benefit varied across models, with some systems improving under knowledge-enhanced few shot prompting and others performing best under simpler strategies. Comparison with state-of-the-art traditional ML baselines showed that, while LLMs do not yet surpass supervised methods, the performance gap has narrowed substantially, particularly on MMD, where the best single-run gap was 2.04 percentage points in F1 and the mean gap under robustness analysis was approximately 5.3 points. Robustness analysis across multiple few-shot example sets confirmed stable performance on WBCD (F1 = 91.37 ± 0.71%) and BCCD (F1 = 87.46 ± 1.91%), while revealing moderate sensitivity on MMD (F1 = 79.62 ± 2.87%). Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models. The results show that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight because clinically significant errors remain.
Habibe Karayiğit, F. Kalelioğlu· International Journal of Int...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.