Jul 2026· JCO Clinical Cancer Informatics· Vol 10 3, pp.
e2500383
· 0 citations· 26 references
Medicine
TL;DR
CTcue performance for patient selection was lower than anticipated, however, it enables accurate, efficient extraction of structured and unstructured EHR data in early-stage NSCLC.
Abstract
Purpose
Manual chart review (MR) of electronic health records (EHRs) is time-consuming, error-prone, and limits the reproducibility and scalability of real-world data (RWD) research. Automation and standardization using natural language processing (NLP) could improve efficiency and scalability. CTcue is an NLP-based software platform designed to extract structured and unstructured data from EHRs. This study evaluated the accuracy and efficiency of CTcue versus MR in patients with early-stage resectable non-small cell lung cancer (NSCLC).
Methods
Included were all patients with stage I to III NSCLC who underwent lung resections between January 2018 and December 2021 at the Leiden University Medical Center, the Netherlands. Demographics, tumor characteristics, treatment, and outcomes were collected. CTcue performance was compared with MR using weighted F1-scores, accuracy, precision, and recall for categorical variables and Bland-Altman analysis for continuous variables.
Results
Eighty-five patients (70.2% of patients from the manual cohort) were identified by both methods and included in the comparison. CTcue achieved weighted F1-scores >0.85 for seven of 15 categorical variables, including sex, tumor location, and deceased status, although some scores were based on low number of observations in both cohorts. Lower performance was observed for variables with varying terminology in documentation, such as Eastern Cooperative Oncology Group status and pathological N-stage. Continuous variables showed negligible mean differences, indicating good agreement. Survival outcomes were identical in both data sets.
Conclusion
CTcue performance for patient selection was lower than anticipated. However, it enables accurate, efficient extraction of structured and unstructured EHR data in early-stage NSCLC. Manual validation remains necessary for variables with varying terminology. Further development of artificial intelligence-based tools-particularly for free-text data extraction-will be crucial to enhance the accuracy and scalability of future RWD research.
This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.
Federica Corso, V. Peppoloni, L. Mazzeo et al.· Communications Medicine· 0 citations
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
PURPOSE
To develop a large language model (LLM) (Truveta Language Model Oncology [TLM-Oncology]) to extract real-world oncology staging data across multiple cancer types from clinical documentation with high precision.
METHODS
We selected patients from a large integrated health system with a bladder, cervical, colorectal, breast, or prostate cancer diagnosis in their structured data. We identified relevant notes using note metadata and keywords and annotated overall stage; T, N, and M; associated timeframe; and cancer diagnosis on a sample of 700 notes as ground truth. Of the 700 notes, 450 were divided equally between training, validation, and test sets for bladder, cervical, and colorectal cancers; 150 were used for targeted error-pattern training on these cancers; and the remaining 100 were split equally between breast and prostate cancer test sets. We started with a pretrained LLM and applied supervised fine-tuning to adapt the model to structured clinical information extraction. Model performance was measured using precision, recall, and F1 scores at the relation level and individual attribute level.
RESULTS
We extracted over 2.5 million staging records for 217,768 patients from over two million notes. Relation-level precision across the six attributes ranged from 0.77 to 1.0 for the first three cancers and, without further training, 0.83 to 1.0 for two additional cancers.
CONCLUSION
TLM-Oncology extracted detailed cancer staging information for five cancers from a variety of clinical documentation within a single integrated health system with high precision and turned data that were previously inaccessible into a valuable resource for downstream use. We are currently evaluating TLM-Oncology on other solid tumors within three additional health systems to assess its generalizability.
S. Abhyankar, Rajesh Rao, Mehraveh Salehi et al.· JCO Clinical Cancer Informat...· 0 citations
Breast cancer remains one of the most common and life threatening cancers worldwide, and early detection is strongly associated with improved survival and reduced treatment burden. This study investigates the ability of Large Language Models to perform diagnostic prediction from structured breast cancer related data. We systematically evaluated 12 LLMs across three public datasets with different clinical characteristics: the Wisconsin Breast Cancer Dataset (WBCD) based on cytological features, the Breast Cancer Coimbra Dataset (BCCD) based on metabolic biomarkers, and the Mammographic Mass Dataset (MMD) based on mammographic attributes. The evaluation covered 13 prompting strategies, including three zero-shot variants, few-shot prompting, and multiple Chain-of-Thought (CoT) and knowledge-enhanced reasoning settings. Performance was assessed using confusion-matrix-based metrics, including accuracy, precision, recall, F1-score, specificity, and Matthews Correlation Coefficient. The results showed that performance was strongly dependent on both dataset type and prompting design, and no single model dominated all tasks. The best model–strategy pairvaried by dataset: Cogito-v1-preview-qwen-32B achieved the highest F1-score on WBCD with an F1-score of 92.00%, GPT 4.1 and GPT 4o on BCCD with an F1-score of 85.39%, and Gemini 2.5 Flash Lite on MMD with an F1-score of 82.91%. Prompt engineering had a substantial effect on outcomes, but its benefit varied across models, with some systems improving under knowledge-enhanced few shot prompting and others performing best under simpler strategies. Comparison with state-of-the-art traditional ML baselines showed that, while LLMs do not yet surpass supervised methods, the performance gap has narrowed substantially, particularly on MMD, where the best single-run gap was 2.04 percentage points in F1 and the mean gap under robustness analysis was approximately 5.3 points. Robustness analysis across multiple few-shot example sets confirmed stable performance on WBCD (F1 = 91.37 ± 0.71%) and BCCD (F1 = 87.46 ± 1.91%), while revealing moderate sensitivity on MMD (F1 = 79.62 ± 2.87%). Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models. The results show that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight because clinically significant errors remain.
Habibe Karayiğit, F. Kalelioğlu· International Journal of Int...· 0 citations
An open-source, human-verified workflow using large language models can accelerate electronic health record abstraction while improving accuracy and supports broader adoption of transparent artificial intelligence methods in clinical research.
Carl Jannes Neuse, Malte Janssen, S. Ibing et al.· BMC Medical Informatics and...· 0 citations
INTRODUCTION
The implementation of imaging features in administrative databases and electronic health records is limited by non-standardized free-text radiology reports. We developed a natural language processing (NLP) tool to automatically extract imaging-based variables from radiological reports and integrate them with clinical data to predict disease course in children and adults with Crohn's disease (CD).
METHODS
Free-text reports from Magnetic-Resonance Enterography and Computed-Tomography Enterography of patients with newly diagnosed CD were linked to clinical data from the nationwide epi-IIRN cohort and processed using Hierarchical Structured Matching Prediction BERT (HSMP-BERT), an NLP model. The primary outcome was difficult-to-treat course, defined by steroid-dependency, the need for ≥2 classes of biologics or surgery. Predictors were identified using Cox proportional hazards and Gradient Boosting Survival Analysis (GBSA) machine learning models.
RESULTS
Among 780 newly diagnosed patients, 160 (20%) developed difficult-to-treat course. Imaging-based predictors, including stricturing/penetrating disease, disease location, disease extent and the MaRIAs score, demonstrated modest discrimination for difficult-to-treat course (0.59 [95%CI 0.53-0.65]). Clinical predictors (laboratory values, induction treatment, age, sex and perianal involvement), showed better discrimination (AUC of 0.68 [95%CI 0.64-0.73]), while combining imaging and clinical variables resulted in only marginal improvement (AUC of 0.69 [95%CI 0.64-0.74]). In GBSA, the AUC was 0.57 (0.52-0.65) for radiologic model, 0.65 (0.62-0.72) to clinical model and 0.67 (0.61-0.72) to the integration model. For surgery, the improvement was more pronounced, in both, Cox regression and GBSA. Across models, the most influential predictors were induction treatment with systemic steroids, and stricturing or penetrating disease.
CONCLUSIONS
Automated NLP-based extraction of imaging reports linked to clinical and laboratory data enables scalable and standardized phenotyping of CD in large datasets and populations.
O. Atia, Gili Kurtser, G. Focht et al.· Scandinavian Journal of Gast...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.