Skip to content
Review

Large Language Models in Medicine: Opportunities, Limitations, and Future Directions

TL;DR

Current evidence indicates that LLMs have substantial potential to enhance healthcare delivery, research, and personalized medicine, but they should currently be regarded as supportive tools rather than autonomous clinical decision-makers.

View source

Similar papers

Review Jul 2026

Large language models in clinical and healthcare scenarios: a global informatics analysis

This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.

Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al. · 0 citations
Review Jul 2026

Large Language Model: Future of Healthcare Research With Challenges

The integration of Large Language Models (LLMs) into healthcare is poised to revolutionize various aspects of medical practice, including clinical decision‐making, patient care, and medical research. This review explores the applications of LLMs such as ChatGPT‐3, ChatGPT‐4, and BERT in healthcare, focusing on their potential to enhance disease diagnosis, treatment planning, and personalized care. The paper presents a comprehensive bibliometric analysis of the growing body of research, highlighting key trends, influential authors, institutions, and geographical contributions. Despite their promise, significant challenges remain, including model accuracy, data privacy, ethical concerns, and the need for domain‐specific fine‐tuning. This review examines the moral and technical challenges associated with deploying LLMs in healthcare, including biases, a lack of transparency, and issues related to model interpretability. The paper further emphasizes the importance of robust frameworks for ensuring ethical usage. It proposes future research directions to address these challenges, including the development of specialized healthcare models, enhanced transparency, and improved integration into clinical workflows. Ultimately, this review aims to inform healthcare professionals, researchers, and policymakers about the transformative potential of LLMs in healthcare while underscoring the critical issues that must be overcome for their widespread adoption. This article is categorized under: Application Areas > Health Care Fundamental Concepts of Data and Knowledge > Big Data Mining Technologies > Artificial Intelligence

Md Belal Bin Heyat, A. Rehman, H. M. Zeeshan et al. · 0 citations
Review Open access Apr 2026

Large language models in hepatology: A systematic review

Background and Aim The rapid advancement of generative artificial intelligence (AI), particularly large language models (LLMs), has opened new frontiers in healthcare, with emerging implications for hepatology. This systematic review synthesizes the current state of research on the application of LLMs in hepatology, focusing on their capabilities in real-world clinical settings, limitations, and future directions. Materials and Methods Electronic databases, including MEDLINE, EMBASE, and OVID as a search platform, were used to identify eligible studies from inception to January 2025. Eligible studies investigated the clinical utility and performance of LLMs in hepatology, with a clear comparison to a defined ground truth. Key findings were extracted and synthesized narratively. The ROBINS-I tool was used to assess the risk of bias in each study. Results Twenty-one studies were included in this review. Our analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease. LLMs effectively assisted with radiological image interpretation, provided clinical decision support, and generated patient education materials. However, the accuracy of these models was highly variable, depending on the specific task and the complexity of the clinical scenario. Limitations, such as the generation of inaccurate or misleading information (“hallucinations”), dependence on training data quality, and ethical considerations, were identified across multiple studies. Conclusion Generative AI demonstrates feasibility across various hepatology applications, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety. Future integration necessitates further research into training methods, data quality, ethical considerations, and real-world validation against standardized benchmarks.

T. Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul et al. · 0 citations
Review Open access 2025

Large Language Models in Healthcare: Opportunities and Ethical Challenges

A thorough review of the developments in LLM technologies, their uses in clinical and administrative settings, as well as their ethical considerations are reviewed to suggest a conceptual structure for responsible implementation that will ensure both technological innovation and patient safety, as well as regulatory compliance and ethical health care practices.

Noah Wright · 0 citations
Review Open access 2026

The Evolution of Clinical Intelligence Through GenAI Co-pilots: A Systematic Review and Thematic Synthesis

Generative artificial intelligence (GenAI), large language models (LLMs), and multimodal foundation models are rapidly transforming healthcare by extending artificial intelligence beyond traditional predictive analytics toward collaborative clinical intelligence. Recent advances have enabled applications in clinical decision support, medical documentation, workflow optimization, patient communication, and personalized care planning. However, the existing literature remains fragmented across technical evaluations, specialty-specific applications, and governance discussions, resulting in a limited understanding of how these technologies collectively function within clinical environments. This study conducted a systematic review and thematic synthesis to examine the emerging role of GenAI as a clinical co-pilot in healthcare. Following the PRISMA 2020 framework, literature was retrieved from Web of Science, PubMed, and Scopus databases. A total of 3,124 records were identified, and 41 studies published between 2023 and March 2026 met the eligibility criteria and were included in the final qualitative synthesis. Quality assessment indicated that 87.8% of the included studies were classified as moderate or high quality, providing a robust methodological basis for the thematic synthesis. Thematic analysis revealed three interrelated domains underlying the evolution of GenAI-enabled clinical intelligence: 1) Perception and Fact Anchoring, involving multimodal data integration, retrieval-augmented generation (RAG), and domain-specific medical intelligence; 2) Clinical Agency and Collaboration, encompassing ambient digital scribing, agentic clinical decision support, workflow augmentation, and personalized care; and 3) Governance and Responsibility, including human-in-the-loop oversight, privacy protection, regulatory compliance, fairness, and trustworthy AI practices. Based on these findings, a conceptual Clinical Co-pilot Framework is proposed to position GenAI as a collaborative partner that supports clinicians rather than replaces them. The framework provides a conceptual basis for future empirical validation and may help inform the responsible implementation of GenAI in healthcare.

Lina Cheng, Chia-Yu Hung, Te-Nien Chien · 0 citations