Skip to content
Book Open access

The Decision Twin: Metaverse Patient Digital Twins as Executable Clinical Reality

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 13157-13163 · 0 citations · 35 references

Abstract

Clinical AI is no longer bottlenecked only by model performance; it is bottlenecked by the accountable interaction loop through which clinicians and patients inspect evidence, test alternatives, and remain responsible for decisions under uncertainty. We argue that today's digital-twin systems, XR interfaces, and foundation-model copilots each address part of this loop, but they fail when deployed as separate products: predictions are not replayable across time, XR becomes descriptive visualization without executable state, and copilots are fluent without auditable grounding. We introduce the Metaverse Patient Digital Twin (MPDT) as a decision-grade clinical artifact defined by one requirement: every displayed claim or simulated scenario must be traceable to a versioned patient state, explicit assumptions, and replayable interaction logs. We specify minimal acceptance criteria, a reference loop in which stakeholders observe new evidence, update the twin state, run bounded simulations, generate explanations, commit decisions, and then log and monitor outcomes, along with a compact architecture that binds interoperability, simulation, governed interaction, and lifecycle controls. Finally, we outline workflows (risk stratification, diagnosis support, treatment rehearsal, training, cross-site coordination) and the evidence required to make MPDTs defensible: calibration over time, subgroup reliability, category-error prevention (observation vs simulation), and measurable workflow outcomes.

Read PDF

Similar papers

Preprint Jul 2026

The Fidelity and Feedback Traps: The Case for Health Digital Twins as Modular Evolving Causal Systems

Digital twins for health may be used to compare treatments, project patient trajectories, and support clinical decisions. While related to mechanical digital twins, those initially developed for engineering applications, replicating the mechanical digital twin architecture and goals may fail in health for two reasons. The fidelity trap is the belief that an accurate model can answer what-if questions by virtue of its accuracy. Prediction and counterfactual reasoning are different tasks, and a twin that can fit past trajectories well may miss the mark when ranking treatments. The feedback trap arises when the twin updates on data its own recommendations helped generate. Refitting in this way can recover a biased relationship and grow more confident even as data grows thinner. We contend that health digital twins should be conceived as causally valid, modular, and evolving systems. Modularity isolates the data and models needed for interventional recommendations, causal validity supports such claims, and governed evolution updates the twin while accounting for how its recommendations reshape the data. We conclude that the standard for a health twin should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.

Nikki L. B. Freeman, Yating Zou, Kyungbok Lee et al. · 0 citations
Review Open access Aug 2026

Validating medical digital twins for clinical decision support: beyond predictive accuracy

Abstract Objective To clarify how validation requirements should be specified for medical digital twins used in clinical decision support, particularly when such systems are intended to compare interventions, treatment timings, dosages, or sequential care strategies. Perspective Medical digital twins are heterogeneous systems that may combine prediction, simulation, mechanistic modeling, machine learning, data assimilation, uncertainty quantification, and decision-support functions. Their evaluation should therefore be driven by their intended use rather than by a single definition of what a digital twin is. For digital twins used primarily for visualization, monitoring, or short-term forecasting, predictive accuracy, calibration, discrimination, and robustness may be the central validation targets. However, when digital twins are used to support intervention-oriented clinical decisions, retrospective accuracy under historical clinical practice is insufficient on its own. Key message Intervention-oriented digital twins address action-conditioned questions: what is predicted to happen under specified alternative actions, assumptions, time horizons, and clinical contexts. Their validation should therefore extend beyond scalar performance metrics to include uncertainty representation, updating stability, robustness under regime change, action-regime validity, counterfactual consistency, clinically weighted error, and decision-level consequences. This requires drawing on established traditions in forecast verification, causal inference, uncertainty quantification, model verification and validation, decision theory, control theory, and post-deployment monitoring. The level of causal or mechanistic support required should match the clinical claim being made, whether at the genotype, phenotype, physiological, or care-process level. Conclusion The scientific-instrument framing is proposed as a pragmatic validation lens for intervention-oriented digital twins, not as a universal definition of digital twins. It helps define the scope within which their outputs can support clinical reasoning. Medical digital twins should be accompanied by explicit validation statements specifying their target population, prediction horizon, supported interventions, uncertainty bounds, and known failure conditions.

Alexandre Vallée · 0 citations
Preprint Aug 2026

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over.

Augusto Bernardo Pissarra, Victor Farias DE Souza · 0 citations
Jul 2026

PRomop: A Decision-Ready Longitudinal Patient Health Record on the OMOP Common Data Model

Objective: Health systems and biopharma face a gap between holding patient data and acting on it: records are fragmented, manually mapped, and structured for storage rather than decisions, so every application re-derives patient state. We present PRomop, an open-source longitudinal record that closes this gap. Materials and Methods: PRomop builds on the OMOP Common Data Model (CDM 5.4) with oncology extensions and adds PatientRecord, a flattened projection collapsing each patient's longitudinal history into a single decision-ready 304-column row. State derivations - lines of therapy, disease status, normalized biomarkers - are computed once at projection time, so analytics, trial matching, and standard-of-care evaluation read one substrate. Results: PRomop is deployed by two oncology organizations - the independently governed HealthTree Foundation (~14,000 patients) and CancerBot (~3,500), a HealthKey-owned deployment - matching against 19,500 recruiting trials across five cancer types. A 20-criterion eligibility search requiring 27-39 joins over raw OMOP reduces to zero against the projection. On a synthetic 1000-patient breast-cancer cohort, eligibility screening averaged 0.30 ms via PatientRecord versus 11.0 ms from raw OMOP, a ~36.8x speedup. Discussion: The projection's significance is as a foundation for other applications: it lowers each one's marginal cost by computing error-prone clinical derivation once and removing it from every consumer. Line-of-therapy inference showed decision-readiness demands embedded clinical reasoning, and that the projection is a living artifact requiring maintenance. Conclusion: A flattened, decision-ready projection over a standards-based longitudinal record is a deployed pattern for turning fragmented data into actionable infrastructure, while remaining OMOP-conformant. Benchmarks measured a ~36.8x eligibility-screening speedup.

A. Blum, Louis Ferger-Andrews, Steven Labkoff · 0 citations
Review Open access Jul 2026

From Algorithm to Bedside: A Clinician's Framework for AI in Practice

This article proposes seven questions that clinicians can run through to evaluate any clinical AI tool in the time it takes to read an abstract, alongside a traffic-light schema for matching oversight to risk and a short list of demands clinicians should make of vendors and institutions.

Alaa Abdelqader, M. Alkhateeb, Abdullah Al-Marrawi et al. · 0 citations
Review Aug 2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.

Veronica Chatrath, Bryan Zhu, George Pu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.