Skip to content
Review Open access

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.

Aug 2026 · Journal of Medical Internet Research · Vol 28, pp. e98184 · 0 citations · 121 references
Medicine

TL;DR

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Abstract

Background

Integration of large language models (LLMs) into health care has accelerated rapidly, yet reliability concerns pose potential risks to patient safety. Although human evaluation has been widely used as an important approach for assessing LLM reliability, a systematic understanding of how such evaluations have been operationalized across studies remains limited.

Objective

This study aimed to characterize the current landscape of human evaluation frameworks for LLM reliability in health care and to identify similarities and differences between the clinical and public health domains.

Methods

In line with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines, PubMed, Web of Science, the Cochrane Library, CINAHL, and Google Scholar were searched for studies published from January 2016 to July 2025. Eligible studies were English-language original research conducted in health care settings that assessed the reliability of LLM-generated responses through human evaluation. Key exclusion criteria were studies without human evaluation and studies focused primarily on LLM model selection, performance optimization, or technical development. Extracted data were analyzed across 3 dimensions: what was evaluated, who evaluated, and how evaluation was conducted. Reported methodological limitations were also categorized and compared between the clinical and public health domains.

Results

Of the 4347 records identified, 71 studies were included in the final analysis (clinical, n=26; public health, n=45). Six reliability indicators were used: accuracy, relevance, completeness, clarity, safety, and consistency. The clinical domain more frequently assessed guideline concordance, internal consistency, and structural coherence, whereas the public health domain more frequently assessed understandability, harm potential, and repeat response consistency. Single-specialty clinicians were the most common evaluators in both domains, although mixed evaluator panels were observed only in the public health domain. Evaluator panels generally consisted of 5 or fewer members. Five-point Likert scales and researcher-defined rubrics were commonly used evaluation approaches in both domains. Key methodological limitations included evaluator subjectivity, nonstandardized indicators, and limited evaluation scope and settings.

Conclusions

To our knowledge, this is the first review to systematically examine how human evaluations of LLM reliability have been conducted across health care. The focus of reliability evaluation differed across domains, with clinical evaluations giving relatively greater attention to clinical validity and logical rigor, whereas public health evaluations gave relatively greater attention to understandability, practical use, and safe use. These differences suggest that the reliability of health care LLMs is difficult to evaluate adequately using a single universal standard. In addition, the methodological limitations identified in this review indicate that current human evaluation approaches are insufficiently standardized. Therefore, future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts and encompass indicator definitions, judgment criteria, evaluator guidance, and evaluation procedures.

Read PDF

Similar papers

Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Review Open access Aug 2026

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Abstract Background Large language models (LLMs) are rapidly emerging in health care, offering opportunities in decision support, education, and research, but raising critical concerns about safety, reliability, and ethics. Although several guidelines for trustworthy AI exist in business and technology, few systematic reviews have applied them to medical contexts. Objective This study aimed to conduct a systematic review of LLM research in health care, applying the AI Guidelines for Business as a framework across 11 domains, including safety, reliability, ethics, transparency, fairness, inclusiveness, privacy, security, robustness, data quality, and verifiability. Methods Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines (retrospectively registered on the Open Science Framework; DOI 10.17605/OSF.IO/P4KSB), the PubMed, Scopus, Web of Science, arXiv, and IEEE Xplore databases were searched on January 15, 2025. Records were screened in 2 stages by 3 reviewers (with records retained only upon unanimous agreement). A total of 247 studies were included, of which 211 (85.4%) contributed quantitative values. Eligible studies were classified across 11 trustworthy AI domains. Heterogeneous metrics were summarized within metric families; when multiple models were evaluated, the mean across models was used as the primary estimate, with best, median, and primary-model sensitivity analyses. The LLM-assisted categorization (GPT-5 mini) was validated by using an automated internal consistency check, and 95% CIs were estimated by using cluster bootstrap on study-level values. Results Of the 25,156 records, 247 (1.0%) studies were included, and of these, 211 (85.4%) contributed quantitative values. Evaluation concentrated on accuracy (143/247, 57.9%) and fairness and inclusiveness (93/247, 37.7%), followed by data quality (47/247, 19.0%) and prevention of misinformation (44/247, 17.8%). Normalized performance was moderate to high (accuracy mean 0.73, 95% CI 0.7-0.76; data quality: 0.64; prevention of misinformation: 0.82). Selecting the best-performing model inflated domain means by up to 0.05. Privacy protection (2/247, 0.8%) and security assurance (0/247, 0.0%) were almost entirely absent. Domain assignments were recoverable from objective metric types in 98.9% of values (Cohen κ=0.985). Conclusions Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability. This finding reflects gaps in reporting rather than demonstrated poor performance, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al. · 0 citations
Review Aug 2026

How large language models can be used for teamwork and communication in healthcare settings: A scoping review.

BACKGROUND Generative artificial intelligence, particularly large language models (LLMs), has rapidly advanced and shows promise in healthcare for supporting teams through their ability to understand and generate medical text. While human-AI collaboration has been explored, the integration of LLMs into healthcare teams remains under-researched. OBJECTIVE This scoping review aims to examine how LLMs are currently used to support teamwork and communication in healthcare teams, including both solely professional teams and those involving patients. METHODS Following PRISMA-ScR guidelines, we registered our review with the Open Science Framework (July 30, 2025). We searched PubMed, Web of Science, and ScienceDirect for articles from 2014 to 2024. After screening 3,865 unique titles and abstracts, 127 full texts were reviewed. RESULTS Twenty studies were included, predominantly employing quantitative and simulation-based designs, with limited in situ evaluations. LLM use cases were categorized into decision support, communication, and administrative functions. Outcome measures primarily focused on accuracy and quality (15/20 studies), with fewer assessing safety (4/20), readability or empathy (5/20), workflow efficiency (3/20), and error modes (2/20). Across use cases, LLMs demonstrated potential to improve efficiency and communication, although performance and risks varied by task complexity and use context. CONCLUSION LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication. However, ethical, legal, and accountability concerns remain. Current studies largely evaluate model performance without considering the dynamics of human team members. Future research should examine LLMs' impact on trust, collaboration, and decision-making within clinical teams, while implementation efforts must address contextual and interdisciplinary factors to ensure responsible integration.

Ilse Super, Olya Rezaeian, Onur Asan · 0 citations
Review Open access Aug 2026

Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals.

OBJECTIVE To evaluate adherence to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) in radiology and medical imaging studies involving large language models (LLMs). MATERIALS AND METHODS We conducted a cross-sectional audit of original LLM research studies published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category. PubMed and Scopus were searched to identify eligible studies. A quota-based subsampling strategy, based on journal publication volume, was used to select approximately 100 studies. All four eligible articles from the Korean Journal of Radiology (KJR) were additionally included as a benchmark. Adherence to the 2025 update of MI-CLEAR-LLM was scored through a two-round, consensus-based process: an initial assessment by one reviewer followed by a critical re-evaluation by secondary reviewers, with consensus adjudication by an additional reviewer when needed. Between-journal differences were analyzed with the Kruskal-Wallis test, followed by Dunn post hoc pairwise comparisons with Holm-adjusted P-values. RESULTS Of 201 eligible studies identified, 102 were finally analyzed after applying the subsampling strategy. Overall adherence to MI-CLEAR-LLM was moderate (mean, 51.2% ± 14.7%; range, 22.2%-84.2%). Adherence was highest for input data type (100%), test-data independence (80.2%), and adaptation strategy (78.1%), and lowest for prompt execution setup (29.4%) and stochasticity management (33.1%). The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%). Adherence varied significantly across journals (P = 0.011), with KJR showing the highest mean adherence (72.8% ± 2.7%). CONCLUSION Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements. Broader adoption of reporting standards is essential to improve the reproducibility and interpretability of future accuracy evaluations.

I. Mese, Saime Turgut Gunes, Ozge Coskun et al. · 0 citations
Review Open access Feb 2026

Methods of Evaluating Large Language Model–Based Health Care Applications Used by Nonprofessionals: Protocol for a Scoping Review

Background Large language models (LLMs) are increasingly used in health care by nonprofessionals (ie, individuals without formal training in health-related professions). These applications must be evaluated in an appropriate manner to prevent misinformation and harmful decisions. To date, guidance to evaluate LLM-based applications for nonprofessional users remains limited and fragmented, leaving researchers and developers without a scientifically grounded set of quality dimensions, metrics, and measurement tools to guide them. Objective This protocol outlines a scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals. It identifies current methods and maps them thematically by assigning them to evaluation dimensions, metrics, and measurement instruments. The review will provide a comprehensive overview of evaluation methods currently in use. Methods The study follows the Joana Briggs Institute approach for conducting scoping reviews and reports. The protocol is reported in accordance with the PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines, and the scoping review will be reported in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. The inclusion criteria comprise studies that evaluate LLM-based applications that are used in the context of health care by nonprofessionals. The search was conducted in PubMed, CINAHL, PsycInfo, and IEEE Xplore. Results since 2021 were considered. Data will be summarized and interpreted qualitatively. Publication screening was conducted by 2 independent reviewers in a blinded manner, with discrepancies settled through discussion. Data extraction and charting will be performed by 1 reviewer. To ensure quality, a random 10% sample of the publications will be independently charted by a second reviewer. Disagreements in the double-extracted subset will be resolved through discussion. Results As of July 2026, a steering committee of 6 researchers has been chosen for the conduct of the review. An initial search resulted in 8538 records after removing duplicates. After screening of these 8538 publications, 17.8% (1524/8538) were eligible for retrieval, of which 88.3% (1345/1524) were retrieved. Full-text screening (completed by 1 reviewer) excluded publications due to nonmatching populations (155/1345, 11.5%), concepts (246/1345, 18.3%), and contexts (24/1345, 1.8%), as well as secondary work (14/1345, 1%), leaving 67.4% (906/1345) of these publications for data extraction. We plan to perform final full-text screening, data extraction, coding, and synthesis of results in the fourth quarter of 2026. Conclusions The scoping review aims to identify and map current evaluation methods for LLM-based applications used in health care by nonprofessionals. It will provide a systematic overview of the current state of research and insights into quality dimensions, metrics, and measurement instruments. The findings will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals. International Registered Report Identifier (IRRID) DERR1-10.2196/93509

Maren Keuchel, Pinar Bisgin, Tom Strube et al. · 0 citations
Review Open access Sep 2026

Comparative performance of large language models for appraising bias in real-world evidence studies.

BACKGROUND Real-world evidence (RWE) is increasingly used to inform regulatory and payer policy decisions and health technology assessment, yet appraising the methodological credibility of RWE studies remains time-intensive and requires specialized expertise. The appraisal task could involve using an appraisal tool that provides a structured approach for evaluating bias in observational studies of comparative effectiveness and safety. Large language models (LLMs) may offer a scalable means to support this appraisal process, but their performance on structured bias assessment tasks has not been fully characterized. OBJECTIVE To compare the performance of LLMs from 6 major artificial intelligence (AI) technology providers against human expert assessments in appraising bias in published RWE studies using the Appraisal of Potential Bias in Real-World Evidence Studies framework. METHODS We conducted a comparative diagnostic accuracy study evaluating 40 LLMs from OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek. Ten published RWE studies representing diverse pharmacoepidemiological designs and data sources were appraised by each LLM using a structured chain-of-thought prompt with conditional rubric injection based on the Appraisal of Potential Bias in Real-World Evidence Studies framework. Two independent human reviewers with pharmacoepidemiology training evaluated each study, with a third adjudicator resolving disagreements to establish the reference standard. LLM performance was assessed using overall accuracy and macro-averaged precision, recall, and F1 scores. Assessment time was compared between models and benchmarked against human reviewers. Bootstrap method was used to construct 95% CI for performance measures. RESULTS Across 280 item-level assessments per model (10 studies × 28 items), overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6 (66.1%), followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). Macro-averaged F1 scores ranged from 30.9% to 66.9%; o3 achieved the highest F1 score (66.9%), followed by GROK-4 (65.5%) and Gemini-3.1-pro-preview (65.4%). Human reviewers required an average of 61.05 minutes per study; all LLMs completed assessments substantially faster, with average time per study ranging from 0.80 to 17.22 minutes relative to humans. CONCLUSIONS LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent. Their greatest value may lie in enhancing efficiency and supporting human-led appraisal as decision-support tools rather than replacing expert review. Future research should assess performance across larger, more diverse RWE study collections and evaluate output reproducibility across repeated runs.

C. Okeke, R. Nechi, J. Khalid et al. · 0 citations