Jul 2026· Journal of Medical Internet Research· Vol 28, pp. e92356-e92356· 0 citations· 41 references
Medicine
TL;DR
This work proposes an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use and demonstrates how the ACF can guide deployment decisions.
Abstract
Abstract Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face practical and ethical challenges related to performance variability, uncertainties about model behaviors, and inequitable performance across languages. Although existing frameworks address domains such as safety, ethics, effectiveness, engagement, and implementation, they often assume or imply—rather than operationalize—an explicit benchmark for deployment and implementation decisions. We propose an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use. The ACF uses project-relevant and off-topic prompts, structured expert review, and prespecified thresholds to produce a documented decision record that can be iteratively rerun after model revisions. We demonstrate the framework through a case application in a tobacco cessation text messaging intervention, illustrating how the ACF can guide deployment decisions.
A pragmatic adoption approach is emphasized to prioritize high-volume, lower-risk clerical and communication workflows; maintain clinician verification and accountability; and apply governance and equity safeguards before scale-up before scale-up.
Michael Christof, Krish Patel, Jian-Dong Zhou et al.· Communications Medicine· 0 citations
A scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals and will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals is outlined.
Maren Keuchel, Pinar Bisgin, Tom Strube et al.· JMIR Research Protocols· 0 citations
Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.
Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al.· Online Journal of Public Hea...· 0 citations
An evaluation framework that pairs the Technology Acceptance Model (TAM) with the Medical Condition Regard Scale (MCRS) and is called the TAM-MCRS LLM Evaluation Framework, a novel clinician-centered approach for comparing which LLMs produce the highest-quality post-care patient education across accuracy, appropriateness, clarity, and completeness.
D. Austria, Grace Williams, Christopher Girardo et al.· JMIR Formative Research· 0 citations
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.
Ilse Super, Olya Rezaeian, Onur Asan· International Journal of Med...· 0 citations
Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.
S. E. McKinney, P. Vu, S. Justice et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.