Skip to content
Review

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

Jul 2026 · arXiv.org · Vol abs/2607.25489 · 0 citations · 107 references
Computer Science

TL;DR

A scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration.

Abstract

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

View source

Similar papers

Review Open access Aug 2026

Multimodal artificial intelligence agents in healthcare: a scoping review

Multimodal artificial intelligence (AI) agents are emerging in healthcare as systems that integrate heterogeneous clinical data, foundation models (FMs), tools, and agentic workflows, but their applications and translational readiness remain unclear. We conducted a scoping review of 37 peer-reviewed studies published between January 2022 and June 2025. Included studies covered clinical decision support ( N  = 17), clinical documentation and report generation ( N  = 3), clinical monitoring and health management ( N  = 13), and medical education and training ( N  = 4). We synthesized modality combinations and fusion strategies, FM utilization and agent architectures, tool integration, agent capabilities, and evaluation practices. Current systems were predominantly text-centric, frequently used closed-source FMs, and remained concentrated in prototype or early technical evaluation stages. Safety, fairness, prospective outcome-based validation, and real-world deployment evidence were limited. These findings suggest that multimodal AI agents are best interpreted as emerging augmentative systems requiring stronger evaluation before clinical translation.

Kai Yu, Shuang Zhou, Yu Hou et al. · 1 citation
Review Open access Aug 2026

Agentic systems in computational pathology: architectures, evidence, and translational challenges

Digital pathology supports whole-slide imaging, remote review, and computational analysis. Most pathology AI systems, however, remain restricted to predefined tasks. Agentic architectures coordinate perception models, language-based reasoning, external tools, and feedback-dependent actions, but their clinical evidence is derived mainly from retrospective benchmarks and research prototypes. We review agentic systems in computational pathology using an operational taxonomy based on dynamic control flow, inference-time tool selection, and knowledge integration. We assess architectures, enabling technologies, and applications in diagnosis, prognosis, and therapeutic support. Reported gains are difficult to attribute to agentic organization because studies differ in backbones, training data, and inference budgets. We therefore emphasize validation scope, computational cost, workflow integration, hallucination and security risks, regulatory requirements, patient preferences, and the conditions under which specialist non-agentic models remain preferable. Agentic architectures have established technical feasibility, but not clinical benefit. Translation should prioritize verifiable tasks, matched comparisons, prospective and external validation, lifecycle governance, and interfaces that preserve pathologist oversight.

Xin-Yu Lu, Qian-Kun Li, Yak Gao et al. · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Open access Jul 2026

Medical AI Agents for Clinical Decision Support: Viewpoint Using the Planning, Action, Reflection, and Memory (PARM) Analytical Lens

This Viewpoint argues that agentic architectures incorporating planning, action, reflection, and memory (PARM) represent a meaningful evolution beyond traditional rule-based, machine learning, and multimodal clinical decision support systems.

Raşit Dinç, N. Ardic · 1 citation
Review Open access Jul 2026

The Next Paradigm in Medical AI: A Survey of Agentic AI in Biomedicine.

This survey provides a structured synthesis of how recent work connects foundation models to governable biomedical agentic systems and distills the recurring challenges and directions identified in the literature for reliable, accountable deployment.

Ali Abdollahi, Mohammad Amin Rezaei, Xi Wang et al. · 1 citation
Review Open access Jul 2026

Agentic AI and Multi-Agent Collaboration in Healthcare: A Comprehensive Survey of Architectures, Clinical Safety, and Future Directions

In the healthcare and clinical domain, artificial intelligence (AI) is evolving from earlier models that primarily predicted outcomes or generated content toward agentic AI systems that demonstrate the capability to make decisions and complete tasks autonomously. Previous research on AI has contributed significantly to disease identification, deep learning applications, large language models (LLMs), and generative AI. These systems primarily function as assistive tools, as they generate text or predictions without directly interacting with clinical infrastructures. Therefore, recent research trends are increasingly oriented toward agentic AI systems that extend beyond traditional predictive and generative model performance. This manuscript provides a detailed review of the current state of agentic AI, starting with the evolution of AI and the concept of a medical agent. A medical agent refers to an intelligent AI system designed to assist in clinical or administrative tasks by analyzing data, supporting decision making, and interacting with healthcare environments. Its underlying agentic AI architecture integrates planning, memory, reasoning, and environmental interaction to enable autonomous tool use, multi-agent collaboration, and continuous perception decision action loops across diverse healthcare applications and clinical workflows. The review further examines safety mechanisms, including human-in-the-loop oversight, self-verification strategies, and regulatory alignment frameworks, which are designed to ensure reliability, accountability, compliance, and safe deployment in regulated healthcare environments. Our findings indicate that a large number of AI agents have been introduced in various manuscripts for healthcare applications; however, fully autonomous systems remain challenging to achieve, as AI still faces several limitations related to reliability, interpretability, data dependency, and integration within complex clinical workflows. In response to these challenges, this survey shifts the focus from task-specific model performance to system-level autonomy and workflow orchestration, providing a structured foundation for understanding the design, deployment, governance, and limitations of agentic AI systems in modern healthcare ecosystems.

Subir Biswas, Rajib Mondal, Manob Saikia · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.