Guideline Concordance of Antibiotic Recommendations From Large Language Models in Dentistry: A Vignette-Based Evaluation in Georgian and English
Abstract
Purpose Antimicrobial resistance is one of the most serious threats to global health, and inappropriate prescribing in primary care is among its principal drivers. Dentistry contributes disproportionately relative to its clinical scope: dental practitioners are estimated to prescribe around one tenth of all commonly used antibiotics, and a substantial share of those prescriptions are issued for conditions in which antibiotics offer no benefit. Professional guidance on this point is unusually consistent. Both the American Dental Association and the Scottish Dental Clinical Effectiveness Programme restrict systemic antibiotics to infections with systemic involvement or evidence of spread, or to situations in which definitive local treatment cannot be delivered promptly. In Georgia the gap between that guidance and practice has been documented on both sides of the clinical encounter. A national survey of 1,876 adults found widespread misconceptions about antibiotics, and a subsequent qualitative study of Georgian dentists found that prescribing decisions were shaped less by guideline content than by diagnostic uncertainty, perceived patient expectation and the absence of an authoritative local reference. That absence is the central feature of the setting: at the time of writing, Georgia has no national guideline on antibiotic prescribing for odontogenic conditions. A clinician who wants an answer must look outside the formal system, and increasingly one of the places they look is a general-purpose conversational artificial intelligence system. These systems are freely accessible, answer instantly in the user's own language and are bound to no jurisdiction. For a clinician without national guidance they occupy precisely the space that guidance would otherwise fill, which makes the accuracy of their advice a practical rather than an academic question. Yet almost all published evaluation of clinical accuracy in these systems has been conducted in English. Performance in languages with a smaller digital footprint has scarcely been examined, and there is reason to expect it may differ, since these models are trained on corpora in which such languages are sparsely represented. The distinction matters clinically, because a dose or a duration that shifts with the language of the question is not a linguistic curiosity but a prescribing error. What this study does Twenty standardised clinical vignettes covering common dental conditions were developed from published guidance, from clinical situations identified as sources of prescribing uncertainty in the investigators' earlier qualitative work, and from the clinical experience of the authors. The set is deliberately balanced so that an antibiotic is not indicated in exactly half of the scenarios. Seven clinical experts rated every scenario for relevance and clarity; all twenty achieved a content validity index of 1.00 on both criteria. One of the seven subsequently joined the study team as a co-investigator; the remaining six had no other role in the study. Each vignette is submitted to three widely available large language models in both English and Georgian, three times each in separate sessions, giving 360 responses. Every response is scored against a reference answer key, frozen before registration and derived independently by two authors from primary guideline documents, on five domains: the decision to prescribe, agent selection, dose and frequency, duration, and referral or red-flag advice. Two raters, blinded to model identity and run number, code every response independently against a coding manual frozen after two calibration rounds. The Georgian version was produced entirely by bilingual dentists using forward and back translation, then pre-tested with seven practising dentists. No machine translation and no language model was used at any stage, either to write the Georgian prompts or to back-translate them. This was a deliberate constraint: using such a system to produce the prompts would align them with the distributional characteristics of the systems under evaluation and could mask the very effect the study is designed to detect. Expected outcomes The primary expectation is that guideline concordance will be lower for Georgian prompts than for English. Concordance is expected to vary across domains, being lowest for dose and duration and highest for the binary decision to prescribe, and responses are not expected to be fully reproducible across repeated runs of an identical prompt. Differences between the three systems are examined as an exploratory question. All four expectations are stated as formal hypotheses in the deposited protocol, with the analysis plan and formulas fixed in advance. The study may also produce findings that are descriptive rather than hypothesis- driven: how often these systems decline to commit to a dose without knowing the user's jurisdiction, how often they name guideline sources they have not consulted, and how often a question asked in Georgian is answered in English. Why it matters If the same clinical question yields different antibiotic advice in Georgian than in English, that disparity falls hardest on exactly the health systems least equipped to detect it: those without national guidance to serve as a corrective. The findings are intended to inform clinicians, dental educators and professional bodies, and to strengthen the case for national prescribing guidance in settings where none currently exists. Materials The protocol, the full scenario set with the reference answer key, all forty prompts in both languages, the coding manual, the pre-randomised run order, the guideline verification record and the pre-testing results are deposited with this registration. Raw model outputs, the completed analysis and the blinding key will be added to the linked project once rating and adjudication are complete.