Skip to content

Agentic AI in Medicine: Challenges for Responsible Development and the Case for Clinical Testing Harnesses.

Sep 2026 · Annals of Biomedical Engineering · 0 citations · 12 references
Medicine

TL;DR

It is argued that each of the four challenges facing responsible development is addressed by a clinical testing harness: a structured evaluation environment comprising scenario libraries built from clinical edge cases, full-trajectory observability, explicit escalation testing, and staged evidence thresholds tied to scope of practice.

View source

Similar papers

Aug 2026

Agentic Foundation Models for Explainable Multimodal Clinical Intelligence: A Human-Centered Framework for Early Disease Diagnosis and Personalized Treatment Planning

A framework that treats diagnosis as a coordination problem rather than a modeling one is described and improved on both across accuracy, F1-score, explanation faithfulness and clinician-rated trust, at an added latency of roughly five seconds per case.

Ganesh Dagadu Puri · 0 citations
Review Open access Sep 2026

State of clinical AI in 2026

Abstract Clinical artificial intelligence (AI) has advanced rapidly, with frontier large language models now matching or exceeding physician performance on simulated diagnostic reasoning and clinical decision-support tasks. Yet adoption has outpaced the evidence base: fewer than 5% of cleared U.S. Food and Drug Adminis...

John Emmett Worth, Anastasia Perez, David Wu et al. · 1 citation
Review Open access Sep 2026

Rethinking AI in clinical decision support: a framework for reciprocal human-AI interaction

Clinical AI systems increasingly match or exceed clinicians on some diagnostic benchmarks. Yet this reveals little about how AI output functions within clinical reasoning or how repeated reliance affects clinicians' independent capability. This paper proposes Bounded Reciprocal Adaptation for Clinician Engagement (BRAC...

C. Greengrass · 0 citations
Preprint Sep 2026

From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning

Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evide...

Min-Ye Shao, Chao-Hui Yu, Yi-Xuan Wu et al. · 0 citations
Review Open access Sep 2026

Artificial intelligence in clinical trials—state of the evidence, gaps, and next steps

Summary Artificial intelligence (AI) affects clinical trials in two distinct but overlapping ways: as the intervention under evaluation and as infrastructure supporting trial design, recruitment, monitoring, endpoint assessment, analysis, and reporting. In this manuscript, we define AI-as-intervention as AI whose outpu...

A. Armoundas, C. Tarabanis, J. Loscalzo · 0 citations
Review Open access Jul 2026

From Algorithm to Bedside: A Clinician's Framework for AI in Practice

This article proposes seven questions that clinicians can run through to evaluate any clinical AI tool in the time it takes to read an abstract, alongside a traffic-light schema for matching oversight to risk and a short list of demands clinicians should make of vendors and institutions.

Alaa Abdelqader, M. Alkhateeb, Abdullah Al-Marrawi et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.