This represents the first attempt to employ a zero-shot, ChatGPT-powered LLM platform to extract RA disease activity measures from real-world EHR data, revealing infrequent real-world documentation.
Abstract
Abstract Objectives We extracted a validated disease activity measure in rheumatoid arthritis (RA), the Clinical Disease Activity Index (CDAI), from a large tertiary academic medical center electronic health record (EHR) using an automated large language model (LLM)-based approach without requiring model pretraining. Materials and Methods The New York Presbyterian/Columbia University Medical Center Clinical Data Warehouse contains EHR data for over 4.5 million patients. RA patients were identified using International Classification of Disease-9 (ICD-9) and ICD-10 codes. Expert-curated CDAI keywords were extracted from unstructured notes using an automated natural language processing (NLP) pipeline leveraging GPT-4o API, a HIPAA-compliant, institutionally approved LLM platform. Performance was evaluated against expert chart review. Results Among 2756 RA patients with notes, 1038 (37.7%) were seropositive, 796 (28.9%) were seronegative, and 922 (33.4%) had unknown serostatus. Clinical Disease Activity Index and its components were extracted in 15.4% (160/1038) of seropositive patients indicating remission or low disease activity. Clinical Disease Activity Index documentation was more frequent among patients with multiple notes and among faculty, with high extraction accuracy (precision/recall/F1 = 0.97). Discussion This represents the first attempt to employ a zero-shot, ChatGPT-powered LLM platform to extract RA disease activity measures from real-world EHR data. Although a low prevalence of documentation was noted, important distinctions were observed when patients were subgrouped by serostatus, level of training, and number of visits. Conclusion An LLM-based pipeline accurately extracted CDAI from a single large academic EHR, revealing infrequent real-world documentation.
Effective detection of ADEs in clinical notes may benefit from NLP models that approximate the clinical reasoning of health care providers as models evolve.
Alan Katz, Abhishek Dhankar, Gillian Fransoo et al.· Journal of Medical Internet...· 0 citations
An open-source, human-verified workflow using large language models can accelerate electronic health record abstraction while improving accuracy and supports broader adoption of transparent artificial intelligence methods in clinical research.
Carl Jannes Neuse, Malte Janssen, S. Ibing et al.· BMC Medical Informatics and...· 0 citations
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
Abstract Background Differentiating among liver disease entities such as autoimmune liver disease (AILD), drug-induced liver injury (DILI), and chronic hepatitis B (CHB) remains clinically challenging due to overlapping clinical manifestations and nonspecific laboratory findings. Conventional machine learning (ML) approaches rely mainly on structured laboratory data, whereas free-text clinical reports and other heterogeneous electronic medical record data are often underused. Large language models (LLMs) may provide a strategy for encoding heterogeneous clinical information, yet their usefulness for liver disease classification remains insufficiently evaluated. Objective This study aimed to evaluate the usefulness of LLM-derived embeddings for clinical data mining in liver disease and to determine whether integrating these embeddings with laboratory variables improves classification across broad disease categories and closely related subtypes. Methods We retrospectively analyzed electronic medical record data from 7543 patients with nonoverlapping liver disease etiologies treated at Beijing Youan Hospital, Capital Medical University, between 2010 and 2025. Three LLMs (Qwen3, Huatuo-o1, and II-Medical) generated semantic embeddings from standardized clinical text, combining free-text examination reports, and structured clinical observations. Performance was assessed in a 3-class etiological task (AILD, DILI, and CHB) and a 4-class task further subclassifying AILD into autoimmune hepatitis and primary biliary cholangitis. We compared embedding-only models, LLM-integrated ML models, and an ML-only baseline using the same structured variable set and preprocessing pipeline, with lightweight natural language processing encoders and zero-shot LLM reasoning as additional comparators. Models were developed using 5-fold cross-validation and evaluated on an internal holdout set using accuracy, macroaveraged precision, recall, and F1-score. Results In the 3-class task, the LLM-integrated ML models achieved macro F1-scores of 0.835‐0.837, compared with 0.791 for the ML-only baseline, with corresponding accuracies of 0.925‐0.929 versus 0.893. In the 4-class task, the LLM-integrated ML models achieved macro F1-scores of 0.717‐0.734, compared with 0.665 for the ML-only baseline, with corresponding accuracies of 0.920‐0.922 versus 0.874. A temporal split sensitivity analysis using cases from 2010 to 2019 for training and cases from 2020 to 2025 for testing showed that the relative advantage of LLM-integrated ML models over the ML-only baseline was preserved. Direct zero-shot LLM reasoning and lightweight natural language processing encoders performed below the embedding-based integrated models. Conclusions In this single-center retrospective cohort of patients with clear-cut, nonoverlapping liver disease etiologies, LLM-derived embeddings provided complementary information to structured laboratory variables for multiclass liver disease classification. The integrated framework showed improved internal validation performance compared with the ML-only model, particularly for non-CHB categories and fine-grained subtype discrimination. Because patients with overlapping liver disease etiologies were excluded, the reported performance may overestimate diagnostic accuracy in broader real-world clinical settings where overlapping syndromes are common. Multicenter external validation and prospective evaluation in more heterogeneous patient populations are needed before clinical implementation.
Unknown authors· Journal of Medical Internet...· 0 citations
Background We present a novel methodological framework for developing and validating a terminology-based Disease Panel to identify and characterize patients with metastatic non-small-cell lung cancer (mNSCLC) in a Spanish cohort by applying clinical natural language processing (cNLP) to electronic health records (EHRs). Materials and methods The mNSCLC Disease Panel was built from standardized vocabularies (Systematized Nomenclature of Medicine–Clinical Terms and Anatomical Therapeutic Chemical) enriched with curated alternative expressions such as synonyms, acronyms, and abbreviations. Terms were organized by clinical relevance and refined through clinical expert annotation. Using EHRead® (Medsavana S.L., Madrid, Spain), a cNLP pipeline, clinical concepts from the Disease Panel were extracted from Spanish EHRs. Performance was validated by clinical experts through the assessment of precision, recall, and F1 scores. Results The Disease Panel comprised 268 terms (mean of alternative expressions 8.4, range 1-41), including 125 (46.6%) demographic and clinical characteristics, 115 (42.9%) treatments, and 28 (10.5%) outcomes. The mean ± standard deviation (SD) precision across all terms was 0.97 ± 0.04, with values of 0.98 ± 0.04 for characteristics, 0.97 ± 0.04 for treatments, and 0.96 ± 0.06 for outcomes. For key terms, the mean ± SD precision, recall, and F1 score were 0.94 ± 0.05, 0.91 ± 0.08, and 0.92 ± 0.05, respectively. Conclusions This is the first validated terminology-based Disease Panel specifically designed for mNSCLC and integrated into a cNLP pipeline. It reliably identifies key clinical features with excellent extraction performance, supporting scalable real-world evidence generation. This approach offers a robust alternative to manual chart review or ‘International Classification of Diseases’ data/claims data in oncology research.
A. Ospina-Serrano, A. Azkarate, S. Ramírez-Peinado et al.· ESMO real world data and dig...· 0 citations
Introduction This study aimed to identify research trends and thematic structures in case reports related to rheumatoid arthritis (RA) and complications published between 1995 and September 2025. By applying unsupervised machine learning (ML), the study sought to uncover longitudinal patterns and cross-domain relationships that may not be fully captured in traditional reviews or large-scale epidemiological studies. Material and methods Bibliographic data were collected from the Web of Science Core Collection using the search term “rheumatoid arthritis and complications,” limited to case reports published between 1995 and 2025. Titles, key words, and abstracts were combined into unified text documents and processed with natural language processing (NLP) techniques. Term Frequency–Inverse Document Frequency was applied for text vectorization. Topic extraction employed non-negative matrix factorization, and clustering was performed using K-means. The optimal number of clusters was determined based on the highest Silhouette Score (0.636). Analyses were conducted in Python (Version 3.10.5) within the PyCharm environment (Version 2022.1.3). Results The unsupervised ML framework identified two major and stable clusters among 1,200 case reports: a pharmacological and immunological management cluster and a surgical and postoperative complications cluster. Cluster 1 included 890 reports characterized by pharmacological and immunological terms such as “TNF,” “therapy,” and “anti,” while cluster 2 contained 310 reports dominated by orthopedic terms such as “arthroplasty,” “knee,” and “revision.” Reproducibility tests across five runs demonstrated consistent clustering patterns, supporting methodological reliability. Conclusions This NLP and ML-based analysis revealed a dual structure in RA complication literature, reflecting immunological treatment-related complications and surgical outcomes. Beyond confirming known domains, the study provides a data-driven characterization of their evolution and overlap, offering insights into the evolution of research themes in RA management. This approach demonstrates the potential of unsupervised ML for longitudinal mapping of clinical research domains.
Naruaki Ogasawara· Rheumatology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.