Skip to content
Conference Open access

A Policy-guided LLM Pipeline for Safer Clinical Decision Support: Error Analysis on High-risk Medication Queries

2026 · SINTEZA · pp. 27-31 · 0 citations · 10 references

TL;DR

This paper presents a policy-guided LLM pipeline for medication-related clinical decision support, designed to classify prompts into safety-oriented decision categories: ACCEPT, WARN, DEFER, ESCALATE, or REFUSE, to handle prompts involving dosage, pregnancy, drug interactions, self-adjustment, and other safety-critical contexts.

Abstract

: Large language models can generate fluent and often convincing answers in medical contexts, but in high-risk settings, fluency alone is not enough. This paper presents a policy-guided LLM pipeline for medication-related clinical decision support, designed to classify prompts into safety-oriented decision categories: ACCEPT, WARN, DEFER, ESCALATE, or REFUSE. The system combines rule-based risk detection, guardrail routing, and a decision policy layer to handle prompts involving dosage, pregnancy, drug interactions, self-adjustment, and other safety-critical contexts. The pipeline was evaluated on a benchmark of 55 medication-related questions using two baseline models. Both models produced the same overall system accuracy of 60.0%, with a false accept rate of 16.4%, suggesting that the main limitations are not model-specific but structural. Three recurring failure modes emerged: false acceptance of implicitly risky prompts, over-escalation of educational or professional-context questions, and under-escalation of self-adjustment or dangerous-intent cases. These findings make the system useful not only as a prototype, but also as a transparent framework for studying where safety-oriented LLM pipelines succeed and where they still fail.

Read PDF

Similar papers

Review Aug 2026

AI-Powered Prescription Error Detection Using Large Language Models (LLMs): A Systematic Review and Future Perspectives

It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.

K. K. Kumar, Koyya Gowtham Reddy, K. Reddy · 0 citations
Open access Jul 2026

General-Purpose vs. Domain-Specific Large Language Models in Antibiotic Clinical Decision-Making: A Double-Blind Evaluation with a 2X2 Factorial Design

Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision, but notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.

Y. Liu, C. Zhang, F. Wang et al. · 0 citations
Open access Aug 2026

A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

DDx-Finder is presented, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns.

H. Lim, H. Yi, J. Yoon et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support

Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format does not reflect how clinical diagnosis is performed in practice. As a result, they remain limited in complex diagnostic settings. We developed Debate-Mixture-of-Agents (DMoA), a novel multi-agent framework that structures role-based interaction to support iterative diagnostic reasoning. Base models and DMoA were evaluated on 297 rare disease cases and 1,719 challenging cases. Across both datasets, DMoA improved most likely diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points over GPT-4o baseline. Ablation experiments showed that the gains were not simply due to the use of more models or longer outputs, but also reflected the contribution of the structured workflow. Further analyses examined how framework design, base model choice, and token budget affected performance. DMoA performed better with a 4*2 structure, stronger base models, and a larger token budget. These findings demonstrate the potential of DMoA for clinical tasks and suggest further investigation of multi-agent frameworks.

Chang Xia, Leilei Ouyang, Huimin Wang et al. · 0 citations
Open access 2026

AARD: A Risk-Aware Agentic AI Framework for Sequential Clinical Decision Support Using Large Language Models

: Modern healthcare systems require decision-support tools that can operate effectively in complex, uncertain, and data-rich environments. This paper proposes the Adaptive Agentic Risk-aware Decision (AARD) framework, a policy-based approach for modelling clinical decision-making as a sequential and adaptive process. The framework integrates structured clinical data, such as vital signs and laboratory measurements, with un-structured clinical text, where Large Language Models (LLMs) are used to extract contextual representations. AARD employs an agentic architecture in which decisions are generated through a learned policy and refined using a risk-aware mechanism. This allows the system to adapt actions over time based on evolving state representations while incorporating safety considerations into the decision process. A recursive learning mechanism enables continuous policy updates using observed state transitions, supporting consistent decision behaviour across sequential steps. The framework is evaluated using open-source healthcare datasets, with performance assessed through both predictive metrics and decision-oriented measures. The results indicate that integrating multimodal representations with risk-aware policy learning provides a structured approach to sequential decision support. Overall, the proposed framework offers a scalable and interpretable approach for combining agentic decision processes, LLM-based feature extraction, and risk-aware optimisation in health-care settings.

G. Jamnal · 0 citations
Jul 2026

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop, and the evidence of safety does not yet exist.

Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.