Skip to content
Open access

Large Language Models for Oncology Guideline Maintenance: Prospective Case Study

Feb 2026 · JMIR AI · Vol 5, pp. e93239-e93239 · 0 citations · 41 references
Medicine

Abstract

Abstract Background Maintenance of oncology clinical practice guidelines (CPGs) is increasingly challenged by the rapid growth of trial data and therapeutic complexity. While large language models (LLMs) have shown promise in information retrieval, their utility in the rigorous, end-to-end workflow of guideline maintenance remains underexplored. Objective This case study aimed to systematically evaluate the performance of frontier LLMs in supporting oncology guideline maintenance. We sought to determine their reliability in predicting necessary guideline updates based on new evidence, their accuracy in extracting data from clinical trials, and their effectiveness as automated auditors for detecting errors in established guidelines. Methods Using the Onkopedia peripheral T-cell lymphoma (PTCL) guideline as a prospective case study, we tasked frontier models with deep-research modes and autonomous web-search capabilities (Gemini 2.5 Pro and GPT o4-mini-high) to predict a guideline update in August 2025 based on the 2021 version. Predictions were validated against the official 2025 revision published in October 2025. Next, we benchmarked evidence extraction accuracy across 80 pivotal trials using models of varying scale (27B-671B parameters vs frontier). Finally, we deployed a stacked LLM workflow to audit 28 recently updated Onkopedia guidelines for linguistic and content-related errors. Results In the predictive task, models captured 36.7% to 40% of substantive updates, often identifying landmark approvals, but frequently overstating evidence. An independent, model-blinded rescoring yielded substantial agreement (weighted Cohen κ=0.75) and confirmed predictive accuracies of 35% to 38.3%. While frontier models demonstrated high accuracy (up to 99.2%) in extracting data from individual studies, substantially outperforming smaller open-source models, this precision declined during multisource synthesis. We observed a position-dependent performance drop in long-form generation, with GPT’s endpoint accuracy dropping from 84.6% in the first half of the drafted guideline to 38.5% in the second half. As automated auditors of existing CPGs, the models successfully identified a median of 16.5 (IQR 13.8-20.3) formal errors per document and detected several clinically relevant inconsistencies (eg, invalid scoring formulas and incorrect staging definitions). Conclusions LLMs currently lack the reasoning stability for autonomous guideline authoring due to deficits in complex synthesis. However, they are effective tools for high-fidelity evidence extraction and automated quality assurance, supporting a human-led, AI-augmented workflow for efficient guideline maintenance.

Read PDF

Similar papers

Open access Jul 2026

Performance of leading large language models in adhering to clinical guidelines for anaplastic thyroid cancer: a comparative study

Leading LLMs show variable capacity to align with ATC clinical guidelines, while top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use.

Mohamed Yasser, Ghada Barakat, S. Awny et al. · 0 citations
Open access Jul 2026

A multidimensional benchmarking framework for large language models in oncologic decision making

A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice.

M. Halıcı, Serkan Saltürk, Irem Sayin et al. · 0 citations
Preprint Jul 2026

Agentic systems for breast cancer treatment recommendations

Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer treatment recommendation generation using 72 real clinical cases across stages I to IV and 1,147 case-specific rubrics generated through Asymmetric Information Rubric Generation (AIRG), in which the rubric generator had access to real clinical decisions unavailable to the evaluated models. Seven pipelines were compared, including single-LLM baselines, tool-augmented systems, and multi-agent architectures with fact checking and autonomous subagent spawning. The best-performing configuration, Claude Opus 4.8 with the D&C+SA pipeline, achieved a global score of 0.594 $\pm$ 0.025. Tool use and increased agent autonomy had mixed effects, improving performance in some settings but degrading it in others. Performance varied by clinical domain and disease stage, and oncologist-led error analysis revealed persistent clinically relevant failures, including incorrect or missing recommendations, flawed justifications, citation errors, outdated claims, and overconfidence. These findings suggest that agentic LLM systems can generate clinically relevant breast cancer recommendations, but remain insufficient for unsupervised clinical use.

Vinicius Anjos de Almeida, N. H. Borges, Leonardo Vicenzi et al. · 0 citations
Review Open access Aug 2026

Benchmarking large language models for question answering on German clinical practice guidelines

The observed performance gains with RAG support the use of LLMs augmented with quality-assured external knowledge in the German healthcare context, while strong performance of open-weight models suggests potential for on-premises deployment in privacy-sensitive clinical environments.

Johannes Schwietering, G. Lichtner · 0 citations
Review Open access Jul 2026

Applications of Large Language Models in Ovarian Cancer Management: Protocol for a Systematic Review and Meta-Analysis

Abstract Background Ovarian cancer (OC) is a highly fatal gynecologic malignancy with complex management challenges and limited long-term survival for advanced stages. Large language models (LLMs)—including systems such as GPT-4, Claude, Google Gemini, and others—are emerging artificial intelligence (AI) tools capable of performing health care–related tasks such as diagnostic support, treatment planning, report generation, and patient communication. However, their applications in OC care have not yet been comprehensively assessed. Objective This protocol outlines a systematic review and meta-analysis aimed at evaluating the use, performance, and clinical impact of LLMs in OC management. We will examine how LLMs have been applied across various domains (eg, diagnosis, prognosis, treatment planning, and patient engagement), the metrics used to assess their performance (eg, accuracy, sensitivity, and area under the curve), and their strengths and limitations. Methods This review will be conducted in accordance with PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines. A comprehensive search strategy will be implemented across biomedical, technical, and Chinese-language databases (eg, PubMed, Embase, Web of Science, IEEE Xplore, and China National Knowledge Infrastructure) from inception to December 31, 2025. Eligible studies include clinical evaluations, validation studies, and real-world implementation reports involving LLMs in OC care. Two independent reviewers will perform screening, data extraction, and quality appraisal using validated tools (eg, version 2 of the Cochrane risk-of-bias tool for randomized trials, Risk of Bias in Nonrandomized Studies of Interventions, Quality Assessment of Diagnostic Accuracy Studies 2, and Prediction Model Study Risk of Bias Assessment Tool+AI). Outcomes of interest include model performance metrics, clinical process impacts, safety concerns, and usability. Meta-analyses will be conducted where feasible using random-effects models in R (meta, metafor, and mada packages), including bivariate models for sensitivity and specificity. Results The review is currently in progress. The PROSPERO registration has been completed, and the literature search and selection process is underway. Study selection, data extraction, and quality assessment are expected to be completed by mid-2026. Final results will include pooled performance metrics (eg, accuracy, F1-score, and area under the curve), qualitative insights into clinical integration, and identification of limitations such as reporting bias or insufficient external validation. Conclusions This systematic review will provide the first comprehensive synthesis of evidence on the application of LLMs in OC care. It will identify promising use cases, highlight safety and reporting challenges, and inform future research directions. The findings are expected to support evidence-based integration of LLMs into gynecologic oncology workflows while promoting transparency and methodological rigor in AI evaluation.

Yanhong Wang, Jialiang Yao, Jianhui Tian et al. · 0 citations
Open access Jul 2026

Assessing clinical decision support system tools in precision oncology: piloting ring testing

Background The EU4Health project PCM4EU aimed to improve survival rates and quality of life of patients with cancer based on precision cancer medicine. To achieve this, enhanced expertise and quality of molecular cancer diagnostics are key. Clinical decision support systems (CDSS) have become increasingly important after the introduction of comprehensive genomic diagnostic profiling. While external quality assessment schemes are mandatory for most diagnostic laboratory tests, similar programs for CDSS tools are currently lacking. To address this, we piloted an international ring test documenting CDSS usage, performance, and manual interpretation. Materials and methods Twenty synthetic datasets were generated, mimicking small variant call sets from a typical targeted 500-gene panel (VCF format) across multiple cancer types (10 tumour-normal pairs and 10 tumour-only). Participants received standardised instructions via e-mail and at a virtual meeting and submitted results using a structured response form. Results Eight laboratories from seven countries participated. All participants submitted results for the 10 tumour-only cases; one submitted results from two assessors. Tumour-only cases contained 6-18 variants where interpretation could be critical. Oncogenic calls for hotspot variants showed good agreement across the various CDSS tools applied; however, variability existed regarding reported variants and clinical interpretation. Conclusion The pilot ring test revealed clinically relevant discrepancies between laboratories and interpreters, underscoring the need for structured external quality assessment schemes for CDSS tools in addition to the existing laboratory workflow schemes. It also highlighted several challenges related to the generation of realistic synthetic data, the design of reporting formats, the definition of ground truth, and the manual interpretation of results.

V. Nygaard, S. Zhao, D. Tamborero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.