Stakeholder role-prompting fundamentally alters clinical decisions and ethical value frameworks of frontier LLMs, with the insurer role producing systematic denial of physician-endorsed, patient-preferred treatments.
Abstract
Background: Large language models (LLMs) are increasingly deployed in healthcare, where they may adopt different stakeholder perspectives, yet the effect of role-prompting on clinical ethical reasoning remains poorly characterized. Methods: We evaluated three frontier LLMs: Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro across 25 ethically complex medical cases. Each model responded from three stakeholder perspectives (physician, patient, insurer) across three independent runs (675 total responses). Decisions were benchmarked against a six-physician panel. Ethical value prioritization was analyzed from physician- and LLM-provided ranked values. A Patient-Centric Decision Index (PCDI) was developed to quantify LLM decision alignment with patient-preferred out-comes. Results: Among 20 cases with clear physician consensus, LLMs prompted as an insurer reduced alignment with physician majority by 50% for GPT-5.4 (p = 0.004), 45% for Gemini 3.1 Pro (p = 0.008) and 10.5% (NS) for Opus 4.6. The insurer role shifted primary ethical values from beneficence (27%) to financial stewardship (20%) across all LLMs. Conclusions: Stakeholder role-prompting fundamentally alters clinical decisions and ethical value frameworks of frontier LLMs, with the insurer role producing systematic denial of physician-endorsed, patient-preferred treatments. These findings raise the need for standardized LLM patient-centricity benchmarks, and physician oversight when LLMs are deployed in clinical decision-making.
This paper investigates whether access to a general-purpose large language model (LLM) improves physicians’ clinical reasoning across diverse healthcare contexts, as well as the possible implications of using an LLM in healthcare settings. Using a randomized controlled trial with 249 physicians in Indonesia, Kenya, and the Netherlands, the study finds that LLM access enhances performance on standardized clinical vignettes in all three countries. The magnitude of improvement varies, with the largest gains observed in Kenya (+18%), followed by Indonesia (+10.7%) and the Netherlands (+7.2%).The results, however, reveal substantial heterogeneity. Performance distributions overlap, and some physicians with LLM access perform worse than those without, indicating that access alone does not guarantee improvement. Higher usage is associated with better outcomes, and less specialized physicians appear to benefit more, implying that LLMs may help reduce skill gaps. Importantly, the findings emphasize that LLMs function as complements rather than substitutes for clinical expertise. However, the study identifies important risks, including automation bias, hallucinations, and context misalignment, underscoring the importance of careful integration, training, and governance.The paper concludes that while LLMs can enhance clinical reasoning, their effectiveness depends critically on how they are implemented within healthcare systems. The paper recommends that policymakers prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone. It also emphasizes the need for investment in infrastructure, continuous monitoring, clear liability frameworks, and inclusive governance to ensure equitable, safe, and context-appropriate deployment. It argues for the importance of social dialogue in managing the process.
N. Rounding, L. S. Arif, Janine Berg et al.· 0 citations
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce CLExEval, a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking. CLExEval combines 5,600 expert-physician annotations with 200 clinical reasoning traces derived from 40 rare diagnostic cases. Our analysis identifies three recurring failure patterns: (i) verbosity bias, where GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity; (ii) a hidden knowledge paradox, where a specialist model reaches 92.5% maximum diagnostic potential but fails to retrieve that knowledge reliably in verbose contexts; and (iii) a 68.6% reasoning-to-output mismatch, where correct diagnoses appear in reasoning traces but are not reflected in final answers. We further evaluate the LLM-as-a-Judge paradigm on a human-verified failure set (n = 142). GPT-4o-mini approved 47.9% of clinically incorrect outputs, while HuatuoGPT-o1 approved all validly scored failures and showed a positive self-preference bias. These results suggest that standalone automated clinical evaluations can substantially overestimate clinical reliability without expert-grounded validation.
Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer et al.· arXiv.org· 0 citations
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its"Hard"subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.39 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 56 of 108 critical (weight-5) criteria (52%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.7% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.
Samiha A. Ismail, Fan X. Chen, Ali Merali· 0 citations
BACKGROUND
Generative artificial intelligence, particularly large language models (LLMs), has rapidly advanced and shows promise in healthcare for supporting teams through their ability to understand and generate medical text. While human-AI collaboration has been explored, the integration of LLMs into healthcare teams remains under-researched.
OBJECTIVE
This scoping review aims to examine how LLMs are currently used to support teamwork and communication in healthcare teams, including both solely professional teams and those involving patients.
METHODS
Following PRISMA-ScR guidelines, we registered our review with the Open Science Framework (July 30, 2025). We searched PubMed, Web of Science, and ScienceDirect for articles from 2014 to 2024. After screening 3,865 unique titles and abstracts, 127 full texts were reviewed.
RESULTS
Twenty studies were included, predominantly employing quantitative and simulation-based designs, with limited in situ evaluations. LLM use cases were categorized into decision support, communication, and administrative functions. Outcome measures primarily focused on accuracy and quality (15/20 studies), with fewer assessing safety (4/20), readability or empathy (5/20), workflow efficiency (3/20), and error modes (2/20). Across use cases, LLMs demonstrated potential to improve efficiency and communication, although performance and risks varied by task complexity and use context.
CONCLUSION
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication. However, ethical, legal, and accountability concerns remain. Current studies largely evaluate model performance without considering the dynamics of human team members. Future research should examine LLMs' impact on trust, collaboration, and decision-making within clinical teams, while implementation efforts must address contextual and interdisciplinary factors to ensure responsible integration.
Ilse Super, Olya Rezaeian, Onur Asan· International Journal of Med...· 0 citations
A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.
Multidisciplinary tumour boards (MDTs) are the standard for gastrointestinal oncological decision-making but remain resource-intensive. Whether specialty-specific role prompting induces genuinely distinct clinical reasoning in large language models (LLMs)—or merely role-appropriate language around an invariant output—has not been systematically tested. We applied five zero-shot prompting frameworks and a majority-vote ensemble to GPT-5 across 100 gastrointestinal oncology cases with MDT-validated decisions: a simulated MDT, multi-expert deliberation, three specialist personas, and a majority-vote ensemble. Concordance with MDT recommendations ranged from 78% to 87%, with no significant inter-framework differences (Cochran’s Q = 8.46, p = 0.133). Specialty-characteristic language was near-universal (97–100%) but uncorrelated with accuracy. Embedding analysis revealed high semantic similarity across personas (cosine similarity 0.805–0.836; η² = 0.049), contrasting with substantially greater output separation under multi-expert deliberation (η² = 0.554–0.581). GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.
Derna Stifini, A. Della Penna, André L. Mihaljevic et al.· npj Digital Medicine· 0 citations