Skip to content
Open access

A clinician-centered evaluation framework for large language models in patient education: Integrating the Technology Acceptance Model and Medical Condition Regard Scale.

May 2026 · JMIR Formative Research · 0 citations · 28 references
Medicine

TL;DR

An evaluation framework that pairs the Technology Acceptance Model (TAM) with the Medical Condition Regard Scale (MCRS) and is called the TAM-MCRS LLM Evaluation Framework, a novel clinician-centered approach for comparing which LLMs produce the highest-quality post-care patient education across accuracy, appropriateness, clarity, and completeness.

Abstract

UNSTRUCTURED Inadequate post-care patient education contributes to preventable readmissions and adverse outcomes that disproportionately affect medically complex, high-need communities. Large language models (LLMs) show promise for generating personalized, plain-language patient education at scale. However, existing LLM evaluation frameworks prioritize technical accuracy over patient accessibility and alignment with health literacy, and few explicitly account for the attitudinal influences that clinician evaluators may introduce into the rating process. In this Viewpoint, we introduce an evaluation framework that pairs the Technology Acceptance Model (TAM) with the Medical Condition Regard Scale (MCRS). We call it the TAM-MCRS LLM Evaluation Framework, a novel clinician-centered approach for comparing which LLMs produce the highest-quality post-care patient education across accuracy, appropriateness, clarity, and completeness. We intend for this framework to be used to evaluate LLM-generated patient education outputs through a two-arm design that pairs an expert clinician panel with automated assessment methods, allowing for inter-arm comparison using clinical vignettes while accounting for measured evaluator attitudinal variance. The framework was developed through the National Institutes of Health Artificial Intelligence/Machine Learning Consortium to Advance Health Equity and Researcher Diversity (AIM-AHEAD) Clinicians Leading Ingenuity IN AI Quality (CLINAQ) fellowship program, in partnership with Ochsner Health and Xavier University of Louisiana. The TAM-MCRS framework integrates two theoretical lenses. TAM maps Perceived Usefulness onto accuracy and completeness, and Perceived Ease of Use onto clarity and appropriateness. Clinicians rate each output with a TAM-based questionnaire, and we then administer the MCRS as a post-scoring attitudinal covariate to see whether their regard for the conditions represented in the vignettes influences those ratings. Together, the two lenses are intended to produce evidence that is objective, theoretically grounded, clinically realistic, and disparity-responsive. Implications for clinician informaticists, health system governance, and responsible artificial intelligence (AI) deployment are discussed. This Viewpoint reflects the authors' position and is written for clinician informaticists, health system AI governance leaders, implementation scientists, and investigators evaluating LLM-generated patient education.

Read PDF

Similar papers

#large language models Review Sep 2026

A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics.

BACKGROUND Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized. OBJECTIVE To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and LLM-generated hospital courses in pediatric inpatient care during early clinical implementation. METHODS We conducted a descriptive evaluation from June 10 to August 8, 2025, at an academic freestanding children's hospital using an Epic EHR with an integrated LLM tool (GPT-4o and GPT-4.1). Clinicians across multiple roles, including attending physicians, residents, and advanced practice providers, reviewed LLM-generated hospital courses for their own patients. Clinicians identified and categorized errors (hallucinations, inaccuracies, or omissions). They also rated text quality (comprehensiveness, conciseness, coherence) on a 5-point scale and perceived harm on an 8-point scale. RESULTS A total of 129 LLM-generated hospital courses were reviewed (median length of stay, 3 days; IQR, 2-7) by 50 involved clinicians. Hallucinations occurred in 21% (95% CI, 14%-29%) of the hospital courses, inaccuracies in 41% (53/129; 95% CI, 33%-50%), and omissions in 24% (31/129; 95% CI, 17%-32%). Overall, perceived harm ratings were low (median, 0; IQR, 0-1). Text quality ratings were high (median [IQR]: comprehensiveness, 4 [3-5]; conciseness, 4 [4-5]; coherence, 4 [4-5]) and comparable with prior literature. CONCLUSION In this pediatric evaluation of LLM-generated hospital courses reviewed by frontline clinicians, errors were common, but perceived potential harm was low, even assuming use without clinician correction. These findings support the use of LLM-generated hospital courses as starting drafts when paired with clinician review and institutional safeguards.

Jasmine E. Kim, Jonathan D. Hron, Daniel J. Kats et al. · 0 citations
Review Open access Jul 2026

An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

This work proposes an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use and demonstrates how the ACF can guide deployment decisions.

Andy J. King, Anthony Banks, L. Hernández et al. · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

Objective: Large language models (LLMs) are increasingly being used to develop medical education materials; however, it remains unclear how reliable, readable, or guideline-compliant the content generated by these models is for non-English-speaking patient groups. We evaluated the quality of Turkish melanoma patient education texts generated by seven frontier LLMs. Methods: A standardized 22-item Turkish prompt, built from international melanoma guidelines, was put to seven models in zero-shot sessions: ChatGPT 4.0 Turbo, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Qwen 2.5 Plus, DeepSeek R1, and Mistral Large 2. Each output was rated for readability (Ateşman Index), understandability, how clearly medical terminology was explained, scientific reliability (DISCERN instrument), empathy, and adherence to a 31-item guideline-based checklist. Model comparisons were summarized descriptively, using model-level absolute scores, score ranges, and rankings. Results: Model performance differed across readability, understandability, reliability, empathy, and guideline-adherence domains. DeepSeek R1 led on both readability (81.6) and understandability (23.5/25). Guideline adherence was strongest for Grok 3 and DeepSeek R1, at 96.8% and 93.5%, respectively, and Grok 3, DeepSeek R1, and Gemini 2.0 Flash each scored above 90% on the normalized total DISCERN measure. DeepSeek R1 also recorded the highest empathy score (90%). Gemini 2.0 Flash had the lowest readability score and produced the longest output (Ateşman 65.8). None of the models provided citations or verifiable sources, so every model received the lowest possible DISCERN Source Reliability score; Mistral Large 2 showed the weakest overall performance. Conclusion: How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next. A few produced text that was clear, empathetic, and reasonably guideline-concordant, but the lack of verifiable citations and uneven guideline coverage remain genuine limitations. These findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Niyazi Çetin, A. Atılan · 0 citations
Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability, usefulness, overall quality, and readability of AI-generated patient education materials. Evidence specifically evaluating AI-generated patient education for hallux valgus, a condition strongly influenced by patient expectations and treatment preferences, remains limited. This cross-sectional comparative study evaluated the performance of three large language models—ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3—in responding to 20 patient-centered questions related to hallux valgus. Questions were developed using AI-assisted question generation and publicly available Google Trends search patterns and categorized into four clinical domains. AI-generated responses were anonymized and independently assessed by three orthopaedic surgeons for reliability, usefulness, and overall quality using 7-point Likert-based reliability and usefulness scales and the Global Quality Scale (GQS). Readability was analyzed using six standardized indices. Inter-rater agreement and between-model comparisons were statistically evaluated. Gemini-2.5-Flash demonstrated modestly higher overall reliability, particularly in questions related to etiology and clinical presentation. DeepSeek-V3 achieved higher usefulness scores in the long-term outcomes and quality-of-life domain and produced significantly more readable content, as reflected by higher Flesch Reading Ease scores and lower grade-level indices. In contrast, Gemini-2.5-Flash generated linguistically more complex responses requiring higher educational levels for comprehension. Overall usefulness and global quality scores did not differ significantly among models. Qualitative review also identified occasional examples of oversimplified or potentially misleading information. Despite these differences, several readability metrics exceeded recommended patient health-literacy thresholds. Contemporary AI-based conversational agents can provide patient-oriented information regarding hallux valgus with variable reliability and readability characteristics, although statistically significant differences were observed across models. A trade-off between factual accuracy and linguistic accessibility was observed. AI tools should therefore be regarded as adjuncts to, rather than replacements for, clinician-led patient education. Awareness of AI limitations and appropriate clinical guidance remain essential to ensure safe, accurate, and patient-centered information delivery. Not applicable.

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations
Jul 2026

A Counsellor-in-the-loop Evaluation Framework for Multi-model Assessment of LLM-generated Mental Health Advisories

The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.

Shahrul Hazman Shamshudeen, N. Sharef, M. S. Yusoff · 0 citations
Review Open access Aug 2026

Understanding primary care physicians’ perceptions of large language model adoption in clinical practice: A qualitative research protocol leveraging the technology adoption behavior framework

Introduction Large language models’ (LLMs’) rapid evolution and intersection with diverse groups and institutions require up-to-date policies, practices, and behaviors to ensure safe and effective implementation. Because physicians play a central role in care provision and face the dual mandate of embracing innovations and safeguarding patient welfare, understanding physicians’ views on LLMs can elucidate the complex interplay of technical, ethical, and professional considerations influencing LLM adoption. Methods and analysis This qualitative study aims to employ a descriptive qualitative design to explore how primary care physicians perceive the adoption of LLMs in the context of their clinical practice. We plan to use semi-structured interviews with purposively sampled primary care physicians from British Columbia, Canada. The data collection will draw on the technology adoption behavior framework, a novel model that integrates the most advanced theories of technological uptake. We expect to use thematic analysis drawing on deductive and inductive approaches to describe physicians’ perceptions. The multidisciplinary research team will prepare and conduct reflexive memos and discussions to ensure nuanced interpretations. Dissemination We aim to disseminate the findings through peer-reviewed journals, professional organizations, and policymaker briefings to support the development of policies, practices, and behaviors that support safe and effective LLM integration into health care. Strengths and limitations The study may provide timely input into relevant policies, practices, and behaviors for policymakers, health professionals, and patients around the use of large language models in health care services. This study uses the technology adoption behavior framework to guide the design of data collection and analysis. The semi-structured interview may reveal the interviewees’ internal views but limits the variety of insights gleaned.

Benny Bikash Pokharel, M. Hsu, Lindsay Hedden et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.