Author

Tom Strube

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Feb 2026

Methods of Evaluating Large Language Model–Based Health Care Applications Used by Nonprofessionals: Protocol for a Scoping Review

Background Large language models (LLMs) are increasingly used in health care by nonprofessionals (ie, individuals without formal training in health-related professions). These applications must be evaluated in an appropriate manner to prevent misinformation and harmful decisions. To date, guidance to evaluate LLM-based applications for nonprofessional users remains limited and fragmented, leaving researchers and developers without a scientifically grounded set of quality dimensions, metrics, and measurement tools to guide them. Objective This protocol outlines a scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals. It identifies current methods and maps them thematically by assigning them to evaluation dimensions, metrics, and measurement instruments. The review will provide a comprehensive overview of evaluation methods currently in use. Methods The study follows the Joana Briggs Institute approach for conducting scoping reviews and reports. The protocol is reported in accordance with the PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines, and the scoping review will be reported in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. The inclusion criteria comprise studies that evaluate LLM-based applications that are used in the context of health care by nonprofessionals. The search was conducted in PubMed, CINAHL, PsycInfo, and IEEE Xplore. Results since 2021 were considered. Data will be summarized and interpreted qualitatively. Publication screening was conducted by 2 independent reviewers in a blinded manner, with discrepancies settled through discussion. Data extraction and charting will be performed by 1 reviewer. To ensure quality, a random 10% sample of the publications will be independently charted by a second reviewer. Disagreements in the double-extracted subset will be resolved through discussion. Results As of July 2026, a steering committee of 6 researchers has been chosen for the conduct of the review. An initial search resulted in 8538 records after removing duplicates. After screening of these 8538 publications, 17.8% (1524/8538) were eligible for retrieval, of which 88.3% (1345/1524) were retrieved. Full-text screening (completed by 1 reviewer) excluded publications due to nonmatching populations (155/1345, 11.5%), concepts (246/1345, 18.3%), and contexts (24/1345, 1.8%), as well as secondary work (14/1345, 1%), leaving 67.4% (906/1345) of these publications for data extraction. We plan to perform final full-text screening, data extraction, coding, and synthesis of results in the fourth quarter of 2026. Conclusions The scoping review aims to identify and map current evaluation methods for LLM-based applications used in health care by nonprofessionals. It will provide a systematic overview of the current state of research and insights into quality dimensions, metrics, and measurement instruments. The findings will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals. International Registered Report Identifier (IRRID) DERR1-10.2196/93509

Maren Keuchel, Pinar Bisgin, Tom Strube et al. · 0 citations