Cloud-Based and Locally Deployed Language Models in Nursing and Health Care: An AI Act–Aligned Framework
TL;DR
Regulatory-aligned LLM integration in pilot university hospitals can enhance health care education and decision-making across standardized taxonomy, evidence-based personalized clinical care algorithms, computational tasks, and nondiscrimination policies, under structured interdisciplinary expert oversight.
Abstract
Abstract Background The integration of large language models (LLMs) into high-risk systems such as health care is accelerating. Rigorous evaluations aligned with emerging legislation are imperative prior to their incorporation into university educational platforms and clinical practice settings. Objective The study aimed at implementing the first Regulation (European Union [EU]) 2024/1689–aligned methodological framework for a systematic, comprehensive, and dynamically adaptable language model evaluation, supporting decision-making in specialized health care management. Methods We analyzed 15 LLMs and 2 small language models. A 7-domain, EU AI Act–aligned methodological framework was used. Feasibility was tested with a dataset of 32 multiparametric-engineered clinical prompts to elicit evaluation in 27 items, with Delphi expert responses as ground truth (available in the repository [32 Clinical Engineered Prompts and Delphi Panel's Responses]). Double-blind interdisciplinary evaluation on a 7-point Likert scale achieved high interrater reliability per model (Krippendorff α=.759 on average). A comprehensive analysis identified specific strengths and vulnerabilities. Safety was analyzed as alignment with both evidence-based nursing and novel structured assessments, including ethical resilience testing via progressive “jailbreaking.” Further novel structured assessments included reference classification, automated consistency, and NANDA-I (North American Nursing Diagnosis Association–International) terminology. Results A stringent “Safety-Gatekeeper” domain immediately classified 11 of 17 language models as unsuitable due to critical failures in evidence-based alignment or ethical resilience. GPT-o1, GPT-4o, Gemini 2.0 Pro Experimental, and 3 Anthropic models surpassed minimum thresholds, permitting evaluation progression. Only Anthropic Sonnet variants achieved uniform “recommended” categorization. For instance, Claude 3.7 Sonnet (extended thinking) produced 75.9% of accurate, focused references, and achieved high average scores both in clinical safety and data security (mean 6.73, SD 0.23 and mean 6.83, SD 0.41, respectively). DeepSeek-R1, Perplexity Sonar, Mistral Large 2, and Qwen2.5-14B-Instruct failed to resist even explicit harmful prompts; Claude 3 Opus resisted both explicit harmful prompts and all jailbreak attempts, while demonstrating null sycophancy. Notably, Qwen2.5-14B-Instruct, operating locally, outperformed 4 of the 15 LLMs in multistep problems in nurse staffing optimization. NANDA-I diagnostic translation capability improved significantly with taxonomy-embedded contexts, with Gemini demonstrating adequate performance (F1-score=0.59, Mean Absolute Priority Distance=4.0). Conclusions Regulatory-aligned LLM integration in pilot university hospitals can enhance health care education and decision-making across standardized taxonomy, evidence-based personalized clinical care algorithms, computational tasks, and nondiscrimination policies, under structured interdisciplinary expert oversight. The methodology demonstrates adaptability across various clinical settings. Future advancements should prioritize multimodal capabilities and locally functioning models, addressing resource disparities in line with Sustainable Development Goal 10, alongside operational resilience, and enhanced data protection.