Patient journey evaluation for consumer AI health assistants
Abstract
Consumer AI health assistants increasingly connect medical records, patient reports, and wearable data across repeated interactions, but evaluation still focuses on isolated prompts and answers. This Perspective has two objectives: to define a journey-level evaluation and governance framework for record-linked patient education, and to specify how longitudinal reliability should be tested before deployment and monitored afterward. We introduce the personal health workspace (PHW), a user-controlled setting for evaluating an assistant as a stateful system across a patient journey. The main outcomes include a clinical journey map and an auditable interaction loop that show what evidence the assistant uses, what actions it takes, and what it carries forward across sessions. Multi-source and long-horizon stress tests use this structure to assess whether claims remain tied to the right evidence, corrections persist, missing or conflicting data are handled safely, and escalation occurs when needed. The framework also specifies deployment requirements for interoperability, privacy, equity, and human oversight. Together, these outcomes support local validation, release decisions, post-deployment monitoring, and prospective studies of clinical effectiveness.