Back to feed
Review Open access

Comparative performance of large language models for appraising bias in real-world evidence studies.

Sep 2026 · Journal of Managed Care & Specialty Pharmacy · Vol 32 9, pp. 1062-1075 · 0 citations · 41 references
Medicine

Abstract

Background

Real-world evidence (RWE) is increasingly used to inform regulatory and payer policy decisions and health technology assessment, yet appraising the methodological credibility of RWE studies remains time-intensive and requires specialized expertise. The appraisal task could involve using an appraisal tool that provides a structured approach for evaluating bias in observational studies of comparative effectiveness and safety. Large language models (LLMs) may offer a scalable means to support this appraisal process, but their performance on structured bias assessment tasks has not been fully characterized.

Objective

To compare the performance of LLMs from 6 major artificial intelligence (AI) technology providers against human expert assessments in appraising bias in published RWE studies using the Appraisal of Potential Bias in Real-World Evidence Studies framework.

Methods

We conducted a comparative diagnostic accuracy study evaluating 40 LLMs from OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek. Ten published RWE studies representing diverse pharmacoepidemiological designs and data sources were appraised by each LLM using a structured chain-of-thought prompt with conditional rubric injection based on the Appraisal of Potential Bias in Real-World Evidence Studies framework. Two independent human reviewers with pharmacoepidemiology training evaluated each study, with a third adjudicator resolving disagreements to establish the reference standard. LLM performance was assessed using overall accuracy and macro-averaged precision, recall, and F1 scores. Assessment time was compared between models and benchmarked against human reviewers. Bootstrap method was used to construct 95% CI for performance measures.

Results

Across 280 item-level assessments per model (10 studies × 28 items), overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6 (66.1%), followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). Macro-averaged F1 scores ranged from 30.9% to 66.9%; o3 achieved the highest F1 score (66.9%), followed by GROK-4 (65.5%) and Gemini-3.1-pro-preview (65.4%). Human reviewers required an average of 61.05 minutes per study; all LLMs completed assessments substantially faster, with average time per study ranging from 0.80 to 17.22 minutes relative to humans.

Conclusions

LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent. Their greatest value may lie in enhancing efficiency and supporting human-led appraisal as decision-support tools rather than replacing expert review. Future research should assess performance across larger, more diverse RWE study collections and evaluate output reproducibility across repeated runs.

Read PDF