2020· International Journal of Machine Learning and Predictive Analytics· 0 citations
Abstract
In recent years, natural language processing has become an important tool in healthcare for extracting useful information from unstructured clinical text such as electronic health records, physician notes, and medical literature. Deep learning has significantly improved the performance of NLP systems, enabling stronger results in tasks such as disease prediction, clinical decision support, and patient risk assessment. However, healthcare NLP still faces major challenges in real-world deployment. Clinical text is often noisy, fragmented, and inconsistent, which can reduce model reliability. In addition, deep learning models lack transparency, which limits their adoption by clinicians who require explainable outputs for clinical decision-making. Privacy and security also remain major barriers because patient data is highly sensitive and subject to strict legal and ethical requirements. Bias in training data can further lead to uneven performance across patient populations. This paper combines a literature review with a healthcare-oriented case study to examine these issues in real-world settings. The findings show that although deep learning offers strong potential for healthcare analytics, progress depends on solving problems related to data quality, interpretability, privacy, and domain adaptation.
A thorough review of the developments in LLM technologies, their uses in clinical and administrative settings, as well as their ethical considerations are reviewed to suggest a conceptual structure for responsible implementation that will ensure both technological innovation and patient safety, as well as regulatory compliance and ethical health care practices.
Noah Wright· International Journal of Mod...· 0 citations
Background and Objective: Early and reliable disease prediction from structured clinical data remains challenging when datasets are small, highly imbalanced, and contain limited positive disease cases. Conventional machine learning (ML) and deep learning approaches often struggle to capture clinically meaningful relationships under such low-data representation conditions due to weak statistical associations between features and prediction targets. This study proposes a clinically grounded GPT2-based table-to-text framework for disease prediction using structured healthcare datasets, motivated by the contextual reasoning capability of GPT models to better capture clinically meaningful relationships when statistical learning alone becomes insufficient due to limited data availability. Methods & Materials: Structured clinical records were transformed into physician-style textual descriptions and enriched through GPT4-generated medical paraphrasing to improve minority-class representation while preserving clinical meaning. Both the original and generated clinical texts were used to fine-tune a GPT2 model across four public healthcare datasets, including heart disease, heart failure, chronic kidney disease, and thyroid cancer recurrence. Gradient-based explainable AI analysis was additionally incorporated to identify clinically important features influencing prediction outcomes. Results: The proposed framework demonstrated consistently strong predictive performance with average precision, specificity, sensitivity, and F1-score of 0.96, 0.97, 0.96, and 0.96, respectively. The model achieved improved sensitivity, stronger generalization, and more stable predictive behavior compared with traditional ML, deep learning, transformer-based, and GAN-augmented approaches. Importantly, the framework consistently emphasized clinically meaningful variables even under severe class imbalance where conventional ML and neural network models often struggled. Conclusions: The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data. By integrating contextual clinical reasoning with explainable prediction mechanisms, the framework demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.
S. Bin Akter, S. Akter, D. Eisenberg et al.· medRxiv· 0 citations
An engineering-oriented, end-to-end roadmap that structures the full lifecycle of clinical language model systems—from model design and domain adaptation to optimization and real-world evaluation is introduced.
The extensive adoption of electronic health records has necessitated the development of automated systems that are capable of understanding unstructured clinical documents. Medical records, such as lab results, radiology findings, and discharge summaries, thus make manual analysis a slow and error-prone process. The paper introduces an AI-driven medical report analysis framework that employs natural language processing and deep learning to automatically locate and interpret the clinically significant information. The system proposed in this paper first preprocesses the medical text to identify the major entities such as diseases, symptoms, and drugs, and then translates them into structured clinical data. An attention-based neural model is used to produce brief analytical summaries, which help clinical decision-making. Experimentally, it was found that the proposed system not only outperformed the manual process in accuracy but also reduced the time. The framework, therefore, increases the efficiency of healthcare and opens up the potential for better utilization of electronic medical records.
Simranjit Singh Bedi, S. Kaswan, Sandeep Singh Kang· International Conference Com...· 0 citations
This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.
Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al.· International Journal of Sur...· 0 citations
The integration of Large Language Models (LLMs) into healthcare is poised to revolutionize various aspects of medical practice, including clinical decision‐making, patient care, and medical research. This review explores the applications of LLMs such as ChatGPT‐3, ChatGPT‐4, and BERT in healthcare, focusing on their potential to enhance disease diagnosis, treatment planning, and personalized care. The paper presents a comprehensive bibliometric analysis of the growing body of research, highlighting key trends, influential authors, institutions, and geographical contributions. Despite their promise, significant challenges remain, including model accuracy, data privacy, ethical concerns, and the need for domain‐specific fine‐tuning. This review examines the moral and technical challenges associated with deploying LLMs in healthcare, including biases, a lack of transparency, and issues related to model interpretability. The paper further emphasizes the importance of robust frameworks for ensuring ethical usage. It proposes future research directions to address these challenges, including the development of specialized healthcare models, enhanced transparency, and improved integration into clinical workflows. Ultimately, this review aims to inform healthcare professionals, researchers, and policymakers about the transformative potential of LLMs in healthcare while underscoring the critical issues that must be overcome for their widespread adoption.
This article is categorized under:
Application Areas > Health Care
Fundamental Concepts of Data and Knowledge > Big Data Mining
Technologies > Artificial Intelligence
Md Belal Bin Heyat, A. Rehman, H. M. Zeeshan et al.· WIREs Data Mining and Knowle...· 0 citations