Skip to content
Review Open access

An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

Jul 2026 · Journal of Medical Internet Research · Vol 28, pp. e92356-e92356 · 0 citations · 41 references
Medicine

TL;DR

This work proposes an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use and demonstrates how the ACF can guide deployment decisions.

Abstract

Abstract Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face practical and ethical challenges related to performance variability, uncertainties about model behaviors, and inequitable performance across languages. Although existing frameworks address domains such as safety, ethics, effectiveness, engagement, and implementation, they often assume or imply—rather than operationalize—an explicit benchmark for deployment and implementation decisions. We propose an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use. The ACF uses project-relevant and off-topic prompts, structured expert review, and prespecified thresholds to produce a documented decision record that can be iteratively rerun after model revisions. We demonstrate the framework through a case application in a tobacco cessation text messaging intervention, illustrating how the ACF can guide deployment decisions.

Read PDF

Similar papers

Review Open access Aug 2026

Evidence, use cases, and implementation safeguards of large language models in primary care

A pragmatic adoption approach is emphasized to prioritize high-volume, lower-risk clerical and communication workflows; maintain clinician verification and accountability; and apply governance and equity safeguards before scale-up before scale-up.

Michael Christof, Krish Patel, Jian-Dong Zhou et al. · 0 citations
Review Open access Feb 2026

Methods of Evaluating Large Language Model–Based Health Care Applications Used by Nonprofessionals: Protocol for a Scoping Review

A scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals and will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals is outlined.

Maren Keuchel, Pinar Bisgin, Tom Strube et al. · 0 citations
Review Open access Aug 2026

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al. · 0 citations
Open access May 2026

A clinician-centered evaluation framework for large language models in patient education: Integrating the Technology Acceptance Model and Medical Condition Regard Scale.

An evaluation framework that pairs the Technology Acceptance Model (TAM) with the Medical Condition Regard Scale (MCRS) and is called the TAM-MCRS LLM Evaluation Framework, a novel clinician-centered approach for comparing which LLMs produce the highest-quality post-care patient education across accuracy, appropriateness, clarity, and completeness.

D. Austria, Grace Williams, Christopher Girardo et al. · 0 citations
Review Aug 2026

How large language models can be used for teamwork and communication in healthcare settings: A scoping review.

LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.

Ilse Super, Olya Rezaeian, Onur Asan · 0 citations
#artificial intelligence Review Sep 2026

A primer on evaluation methods for large language models in healthcare

Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.

S. E. McKinney, P. Vu, S. Justice et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.