Large language models for population-level public health communication: an evidence map of deployment, reach and policy gaps
Abstract
Preprint version 6 (posted 4 September 2026). This version supersedes versions 1–5 and should be used in preference to them. Versions 1–5 were posted on 6 April, 3 June and 4 June 2026. What changed, and why. Version 6 reports a corrected re-analysis carried out after peer review. An audit of the data-charting pipeline found it defective in three respects. Six keywords were matched as unanchored substrings, most consequentially “production” matching inside “reproduction” in the Creative Commons licence text carried by open-access PDFs, so studies were classified as deployed at scale on the strength of licence boilerplate. Two variables were assigned a positive default when no term was found, converting missing data into apparent observations. And the original reliability check compared the classifier against a superset of itself, so it measured agreement between two variants of one instrument rather than accuracy. A human reference standard was therefore established: 50 studies charted by hand, blind to all machine output, from exactly the text the pipeline received. Each variable was reassigned to whichever instrument measured most accurate against that standard, and every affected number was re-derived. Headline changes: real-world deployment 9.6% to 1.1%; user testing 7.2% to 8.5%; English as a target language 83.3% to 12.3%, with 72.1% of studies naming no target language at all; misinformation detection 186 to 235 studies and correction 11 to 21; and the full-text charting substrate corrected from 194 to 171 of 552 (31.0%). Every change is itemised in Supplementary Table S16. The geographic figures are unchanged. Two of these changes strengthen the paper's claims rather than weakening them. Real-world deployment is rarer than originally reported, and the linguistic finding is sharper: the problem is not that the field targets English, but that it largely does not state which language its tools are meant to serve. Large language models (LLMs) are increasingly proposed for population-level public health communication, yet whether this research is translating into deployed tools that reach the populations and policy settings that most need them is unknown. We conducted an automated, computationally assisted evidence map (reported using the PRISMA extension for scoping reviews), screening 30 715 records from six databases and grey literature (2019–2026). AI-assisted screening agreed closely with a second screen (Cohen's κ=0.83). Charting was assigned per variable to whichever instrument measured most accurate against a 50-study human reference standard, giving 56% to 98% exact-match accuracy by variable; full text was available for only 171 of the 552 included studies (31.0%), the rest from title and abstract. Findings resting on the least accurately charted variables are therefore reported as exploratory. The field grew from 3 studies in 2020 to 215 in 2025, but has not translated into practice: 84.1% of studies remained at concept or laboratory stage, 8.5% reported user testing and 1.1% real-world deployment. It is also skewed away from global need. Among the 474 studies with identifiable affiliations, first authorship concentrated in the United States (38%) and other high-income countries; 72.1% never specified a target language, and English predominated among those that did. Both magnitudes are entangled with the review's English-language inclusion criterion. Misinformation detection (235 studies) exceeded correction (21) by an order of magnitude. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation, alongside deliberate investment in multilingual, LMIC-relevant and correction-oriented applications. The analysis apparatus is openly available on the Open Science Framework project node https://osf.io/n8q4y/ : both charting datasets (as originally submitted and corrected), the reconciliation between them, the human reference standard, the accuracy computation and the full pipeline. A defect in the previously deposited dataset, a silently lost first-author-country column, has also been corrected there. The protocol registration is separate, at https://doi.org/10.17605/OSF.IO/X5N78 , and contains the protocol only.