Aug 2026· International Journal For Multidisciplinary Research· 0 citations· 35 references
TL;DR
ClinicalAgentOps is proposed, a framework that relocates governance from the artifact into the agent’s execution path and contributes an explicit argument from the premises of design-time assurance to the necessity of inpath control.
Abstract
Generative AI is entering clinical practice not as a predictor but as an actor. Contemporary healthcare deployments increasingly involve agents: language-model systems that plan over multiple steps,retrieve patient context, invoke tools, write to the electronic health record, and coordinate with otheragents. Health AI governance, however, remains overwhelmingly design-time. Premarket review,transparency labels, and reporting standards evaluate a model artifact under the assumption that behavior is a stable property of that artifact. Agentic systems violate this assumption: their effective behavioris constituted at runtime by the composition of instructions, retrieved context, tool affordances, mem
ory, and inter-agent interaction, none of which is fixed at approval time. This paper proposes ClinicalAgentOps, a framework that relocates governance from the artifact into the agent’s execution path. Itcontributes an explicit argument from the premises of design-time assurance to the necessity of inpath control, stated with its falsifying conditions; a runtime failure taxonomy for clinical agents inwhich the unit of analysis is the action trajectory rather than the input–output pair; a two-dimensionalmodel treating autonomy as a graduated, revocable, per-action grant indexed by a Clinical Action RiskTier, with a stated derivation rule from which the minimum control set for each (tier, autonomy) pairfollows; a five-plane reference architecture spanning authorization, execution, observation, assurance,and accountability; a clinical agent trace schema extending emerging generative-AI telemetry conventions with attribution, evidence, and oversight attributes, together with governance metrics computablefrom it; and a mapping from framework components to obligations under prevailing risk-management,
privacy, and medical-device regimes. This is a framework and position paper: claims about controlefficacy are advanced as falsifiable hypotheses with the study designs that would test them, not asresults
Agentic artificial intelligence, in which a model reasons, calls tools, and acts in a closed loop rather than emitting a single prediction, is arriving in medicine, where the tolerance for confident error is close to zero. This paper takes a design stance on healthcare agentic AI: rather than propose a new model, it asks how to build a clinical agentic system so that its autonomy is bounded by verification and human oversight. We distil six design requirements from the clinical AI literature, covering harm minimisation, human oversight, calibrated uncertainty, verification and provenance, equity, and interpretability. We give a reference architecture in which a reasoning agent is wrapped by a verification gate and a clinician escalation path, and we formalise that gate. We prove two guarantees: escalation is monotone in safety, so widening it never increases expected harm, and a risk-weighted gate minimises expected harm for any fixed amount of clinician effort. We validate the design with a fully reproducible simulation of an agentic clinical triage loop. The escalation gate trades autonomy for safety along a smooth curve, halving expected harm as oversight rises; layering an independent verifier and human review cuts the unsafe-action rate from 0.230 with the agent alone to 0.077; and risk-weighted escalation attains lower harm than confidence-only escalation at every clinician budget, exactly as the theory predicts. The results argue that in medicine the safety of an agentic system is engineered at the level of the loop, through verification, risk-weighted escalation, and clinician oversight, and can be designed and reasoned about rather than left to chance.
Setu S. M. Kodi, Sudheer Singamsetty· International Journal of Com...· 0 citations
This Viewpoint argues that agentic architectures incorporating planning, action, reflection, and memory (PARM) represent a meaningful evolution beyond traditional rule-based, machine learning, and multimodal clinical decision support systems.
Raşit Dinç, N. Ardic· JMIR Medical Informatics· 1 citation
The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.
Rakesh Sharma, S. Pugh, C. Beeche et al.· 0 citations
The findings show that meaningful oversight is not a single human approval step and is a lifecycle capability that combines bounded autonomy, evidence-based escalation, stop authority, continuous validation, audit records, and institutional learning.
Aaron K Montgomery, Hannah E Gallagher, Derrick L Mercer· International Journal of Eng...· 0 citations
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.
Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou et al.· arXiv.org· 0 citations