Skip to content
Review Open access

Large Language Models in Adverse Drug Reaction Detection and Pharmacovigilance: A Systematic Review of Current Applications, Challenges, and Future Directions

Aug 2026 · Diagnostics · Vol 16, pp. 2435 · 0 citations · 76 references
Medicine

TL;DR

Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making, highlighting their potential for precision medicine and big data-enabled safety monitoring.

Abstract

Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.

Read PDF

Similar papers

Review Jul 2026

Performance of large language models in data extraction for evidence synthesis: A systematic review

OBJECTIVE Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to comprehensively evaluate LLM accuracy, reliability, and efficiency for data extraction in evidence synthesis, and to identify optimal implementation strategies. METHODS We searched PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025. Studies were eligible if they evaluated one or more LLMs for data extraction against a human reference standard and reported quantitative performance metrics. Two reviewers independently extracted data and assessed methodological quality using PROBAST + AI and reporting completeness using TRIPOD-LLM. Narrative synthesis was performed due to substantial heterogeneity precluding meta-analysis. RESULTS Twenty-seven studies met inclusion criteria, evaluating models including GPT-4/4o (n = 15), Claude versions 2-3.5 (n = 10), Gemini (n = 3), and open-source alternatives including Llama, Mistral, Qwen, and DeepSeek. Overall accuracy ranged from 47% to 99.9%, with substantial heterogeneity by task type and data granularity. Categorical and string variables were extracted more reliably (74-96%) than numerical data (47-88%). Claude 3.5 Sonnet achieved high accuracy in an assistive workflow (91.0%; 95% CI: 90.4-91.6%), exceeding human-only extraction (89.0%). Claude models outperformed GPT in head-to-head comparisons (OR 1.70 for event counts). Omissions were the dominant error type (60-74%), with hallucination rates of only 0.08-6%, challenging widespread fabrication concerns. Time savings of 33% to 87% were reported, although most included studies did not quantitatively assess efficiency. Methodological quality was generally robust, with 74.1% of studies rated low risk of bias under PROBAST + AI. Mean TRIPOD-LLM compliance was 88.5%, though gaps in inference settings and model version documentation were common. CONCLUSION LLMs demonstrate promising but variable performance for data extraction in evidence synthesis. Current evidence supports their integration as assistive tools within dual-extraction workflows requiring human verification, rather than as autonomous extractors. Categorical data is extracted more reliably than numerical outcomes, and few-shot prompting with structured output formats consistently improves performance. Standardised benchmarks and prospective comparative studies remain priorities for future research.

Ravi Shankar, Amaevia Lim, Xu Qian · 0 citations
Review Open access Aug 2026

Comment on: “A comprehensive landscape of AI applications in broad-spectrum drug interaction prediction: a systematic review” (Marzouk et al., 2025)

Marzouk et al. reviewed 147 studies on artificial intelligence (AI) applications for predicting drug–drug, drug–disease, and drug–nutrient interactions, providing a broad overview of current machine learning and deep-learning approaches. However, several methodological and conceptual limitations reduce the reproducibility and interpretability of the review. The search strategy appears largely restricted to PubMed with title- and abstract-level filtering, while manual record removal is reported without explicit criteria defining “irrelevant” studies, limiting transparency and reproducibility. Protocol registration, duplicate independent screening, standardized extraction procedures, and formal bias assessment using established frameworks such as ROBIS, PROBAST+AI, and TRIPOD+AI were not clearly reported. The review reports performance metrics such as area under the receiver operating characteristic curve (AUROC), but does not provide a structured framework for interpreting or comparing metrics across heterogeneous datasets, prediction tasks, and evaluation protocols. Because the interpretation of AUROC and precision–recall metrics depends on class prevalence, outcome definition, and the intended prediction task, future reviews should report complementary discrimination metrics, calibration, uncertainty estimates, and external validation rather than assuming that any single metric is universally preferable. Claims of superior model performance should be supported by confidence intervals and statistical comparisons appropriate to the evaluation design, such as paired DeLong testing when applicable. Claims of superior model performance should also be supported by appropriate statistical testing, including methods such as the nonparametric DeLong test. Several conceptual clarifications are also warranted. AI models may prioritize hypotheses but do not replace experimental or clinical validation under current regulatory standards. Furthermore, AUROC should not be conflated with pharmacokinetic area under the curve, and SciBERT should not be characterized as a three-dimensional molecular graph framework. Future reviews should adopt transparent multi-database searches, structured bias assessment, and reproducible reporting practices.

Alireza Kargar, Mohammad Ali Zamani, Ghader Mohammadnezhad · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Review Aug 2026

AI-Powered Prescription Error Detection Using Large Language Models (LLMs): A Systematic Review and Future Perspectives

It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.

K. K. Kumar, Koyya Gowtham Reddy, K. Reddy · 0 citations
Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists

LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

Abstract Background Systematic literature reviews (SLRs) are essential for evidence synthesis in health research but remain labor-intensive, especially at the screening stage. Manual review of titles and abstracts requires substantial human effort, while existing automation tools still have limited adoption in health technology assessment. The EQ-5D questionnaire, a widely used patient-reported outcome measure for health-related quality of life, provides data that frequently underpin reimbursement and policy decisions. Objective This pilot study evaluated whether recent large language models (LLMs) can support the identification of publications reporting EQ-5D data in PubMed records, using only publicly available metadata (title, abstract, and keywords). Methods A total of 200 publications retrieved through the EuroQol PubMed filter were manually labeled by experts as reporting or not reporting EQ-5D data. The dataset was split into stratified training, validation, and test subsets. Several machine learning approaches were compared, including a Naïve Bayes baseline using bag-of-words features, a decision-tree model based on full-text keyword occurrence, and transformer-based LLMs (Bidirectional Encoder Representations from Transformers [BERT], Biomedical BERT [BioBERT], Scientific BERT [SciBERT], and Biomedical Language Understanding Evaluation BERT [BlueBERT]). Both classifier-only and fine-tuned configurations were tested across multiple learning rates. Model performance was assessed using accuracy, precision, recall, and F1-score. Results Baseline approaches achieved near-random test performance (accuracy around 0.53). Classifier-only LLMs modestly improved results (accuracy up to 0.64 with SciBERT). Fine-tuned models substantially outperformed these baselines, with BERT and BioBERT achieving the best performance (accuracy=0.70; F1-score=0.68). In screening-oriented evaluation, this configuration achieved 90.0% sensitivity, 40.0% specificity, and 6 false negatives on the held-out test set. The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection. Conclusions This study provides the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature. The findings support technical feasibility but do not establish a reliable stand-alone automated screening tool. Although limited by dataset size, the proposed workflow is reproducible and adaptable to other patient-reported outcome measures. Because validation was based on a single small train-validation-test split, the results should be interpreted as preliminary; future work will scale data collection, include statistical testing, and explore semisupervised learning to further reduce manual screening workload.

Gábor Kertész, J. Czere, Zsombor Zrubka et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.