Skip to content
Review

Safety and security of large language models in healthcare

Aug 2026 · Nature · Vol 656, pp. 577 - 589 · 0 citations · 189 references
Medicine

TL;DR

The Review examines rapid LLM adoption in clinical care, outlining emerging security and safety risks across development stages, key protective layers, clinically relevant threats and current mitigation responsibilities in a single integrated framework.

View source

Similar papers

Review Open access 2025

Large Language Models in Healthcare: Opportunities and Ethical Challenges

A thorough review of the developments in LLM technologies, their uses in clinical and administrative settings, as well as their ethical considerations are reviewed to suggest a conceptual structure for responsible implementation that will ensure both technological innovation and patient safety, as well as regulatory compliance and ethical health care practices.

Noah Wright · 0 citations
Review Jul 2026

Large language models in clinical and healthcare scenarios: a global informatics analysis

This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.

Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al. · 0 citations
Review Open access Jul 2026

Evaluating the safety of large language models in healthcare and dentistry: adversarial testing approaches

The emergence of large language models (LLMs) provides new avenues for clinical support in healthcare and dentistry. However, these models often exhibit unpredictable behaviours when challenged by adversarial or misleading inputs. Recent data indicate that nearly 20% of LLM outputs contain safety risks or biases, necessitating rigorous evaluation prior to clinical use. This review examines AI red teaming, a systematic approach for identifying system vulnerabilities through simulated attacks. It details methodological approaches and outcome measures while proposing a structured framework to integrate these safety evaluations into the clinical AI lifecycle. This review focuses on prompt-based attacks, such as prompt injection and jailbreaking, which are highly relevant in medical settings. It evaluates various testing strategies, including manual expert reviews, automated “attacker” models, and hybrid human-in-the-loop systems. A lifecycle-based framework is introduced, utilizing the collaborative “red-blue-purple” teaming model. This approach spans pre-deployment testing, live deployment monitoring, and iterative review audits to ensure that clinical guardrails remain robust against evolving adversarial tactics. Safe implementation of LLMs in dentistry and healthcare requires continuous, iterative adversarial testing rather than static assessments. Success depends on standardized protocols, multidisciplinary collaboration between clinicians and AI researchers, and the development of domain-specific benchmarks. Bridging existing regulatory gaps through these structured frameworks is vital for ensuring LLMs are safe, reliable, and clinically fit for patient care.

F. Umer, Muhammad Muthar Shaikh, Absar Ur Rahman · 1 citation

Layered security and trust mechanisms for GenAI applications in critical domains

Large Language Models (LLMs) are increasingly being deployed in critical domains such as healthcare, finance, and public infrastructure to support intelligent decision-making and conversational interactions. However, these systems introduce significant challenges related to security, reliability, and trustworthiness. Vulnerabilities such as adversarial prompt injections, behavioral manipulation, and multi-stage attacks can lead to unsafe outputs, privacy risks, and loss of user trust. There is a need for robust approaches that ensure both safe application-level interactions and adaptive system-level defenses against evolving LLM threats. In this thesis, we propose a unified two-layer approach to enhancing the trustworthiness and security of LLM-enabled systems. At the application layer, we develop EmpathAI, a RAG-based mental healthcare chatbot that incorporates source tagging, sentiment-aware context retrieval, and a two-layer defense mechanism using regex filtering and prompt engineering to mitigate prompt injection attacks. Building on this, at the system layer, we introduce the Adaptive LLM Threat Response (ALTR) framework, which integrates behavioral anomaly detection, context-aware prompt classification, and temporal threat memory to identify and mitigate adversarial interactions in real time. We evaluate both layers using conversational datasets and adversarial interaction traces. At the application layer, EmpathAI achieves high semantic alignment (similarity scores >0.80–0.85), with all prompt injection classes successfully mitigated. At the system layer, ALTR attains strong detection performance (accuracy and AUC of 0.961, false-negative rate of 0.9 percent) under low-latency constraints (<20 ms). Together, these results demonstrate that securing LLMs in critical domains requires both application-layer trust and systemlayer defense, and that combining domain-aware RAG systems with adaptive multi-layer security frameworks enables their trustworthy and reliable deployment in high-risk environments.

Vani Seth · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

The rapid growth of artificial intelligence systems (AI systems) has increased interest in the use of patient care and clinical decision-making processes. There is some uncertainty regarding their reliability and safety in clinical practice. A more detailed systematic review of literature examining LLMs applied to healthcare diagnosis was conducted. A PRISMA-based systematic review has been carried out of relevant literature published in the major databases for the years 2022–2025. Key findings include a growing trend to develop multimodal models based on diverse input modalities, combining LLM models with other models as part of clinical workflows. The Usage of complementary methodologies such as retrieval-augmented generation, knowledge graphs, and federated learning is highly expanding, particularly in enhancing the efficiency and accuracy of clinical decision-making processes. Significant challenges such as hallucinations, bias, prompt sensitivity, limited explainability, and inadequate clinical validation continue to pose major obstacles. Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis. Overall, this review contains multiple recommendations for future research in many areas (e.g., LLMs) to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Review

Large Language Models in Medicine: Opportunities, Limitations, and Future Directions

Current evidence indicates that LLMs have substantial potential to enhance healthcare delivery, research, and personalized medicine, but they should currently be regarded as supportive tools rather than autonomous clinical decision-makers.

Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al. · 0 citations