Skip to content
Review Open access

Development of an AI-driven chatbot for medication-assisted treatment standards in Scotland

Jul 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 55 references
Medicine

TL;DR

This study provides the first prototype of an AI-driven chatbot specifically designed for MAT professionals, demonstrating feasibility of integrating advanced AI technologies to address information access barriers in addiction treatment.

Abstract

Background Scotland faces a severe public health crisis with drug-related deaths reaching 267 per million people, ranking second globally after the United States. Medication-Assisted Treatment (MAT) represents a proven intervention for heroin addiction. However, healthcare professionals struggle with accessing and interpreting current MAT standards through fragmented information systems and time-consuming manual searches across multiple websites. Despite advances in healthcare chatbots leveraging Large Language Models (LLMs), no specialized systems exist to support MAT delivery or integrate advanced technologies like Retrieval-Augmented Generation (RAG) and Knowledge Graphs for addiction treatment. Objective To develop and evaluate an AI-driven chatbot prototype that integrates LLMs, RAG, and Knowledge Graphs to enhance healthcare professionals’ access to MAT standards in Scotland, addressing current barriers in information delivery and clinical decision-making. Methods We employed a mixed-methods approach combining a survey of 39 MAT healthcare professionals (31% response rate) and systematic literature review following PRISMA guidelines. The chatbot prototype was developed using Llama2 language model, Neo4j knowledge graphs, and custom RAG implementation. Data was ethically collected from Public Health Scotland and Healthcare Improvement Scotland websites. Performance was evaluated using BLEU and ROUGE metrics, with prototype deployment via Streamlit interface. Results Survey findings revealed significant challenges with current communication methods: only 5 of 39 respondents rated existing systems as “exceptional,” while 17 rated them as “average” or below. Primary challenges included decentralized information (n=13) and time-consuming access processes (n=8). Literature review of 14 healthcare chatbot studies identified a critical gap in MAT-specific applications. The developed prototype demonstrated moderate performance with BLEU score of 36.64, ROUGE-1 score of 0.48, and ROUGE-L score of 0.42. The knowledge graph successfully integrated 227 nodes, 136 relationships, and 8 characteristics representing comprehensive MAT standards. The system successfully retrieved relevant MAT standards information in response to queries about specific MAT standards, medication protocols, and implementation guidance Conclusions To our knowledge, this study provides the first prototype of an AI-driven chatbot specifically designed for MAT professionals, demonstrating feasibility of integrating advanced AI technologies to address information access barriers in addiction treatment. While performance metrics indicate potential for enhancing MAT information delivery, further development is needed to improve semantic understanding and response naturalness. The prototype establishes a foundation for future integration with electronic health records and broader healthcare systems, with the potential to support improved treatment outcomes for individuals with heroin addiction in Scotland, subject to longitudinal clinical validation.

Read PDF

Similar papers

Review Open access Aug 2026

Multilingual Conversational AI Chatbots for Efficient Healthcare Delivery During Case History-Taking: A Systematic Review

Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias.

R. Sharanesha, Deepti Virupakshappa, A. Abushanan et al. · 0 citations
Jul 2026

Healthmate: An AI Chatbot for Symptom-Based Healthcare Consultation

HealthMate is presented, an intelligent, explainable AI chatbot framework designed for preliminary healthcare consultation that demonstrates rapid retrieval, robust natural language comprehension, and clear explainability without replacing professional medical diagnosis.

K. Jyothi, Shaik Khasim Basha · 0 citations
Open access Jul 2026

A PRACTICAL FRAMEWORK FOR EVALUATING AI CHATBOTS IN NON-ENGLISH DRUG INFORMATION SERVICES

Artificial intelligence, particularly AI-based chatbots, has gained increasing attention as a tool to support pharmacists in providing drug information services, especially in settings with pharmacist shortages. This study compared the competencies of ChatGPT, ChatGPT Pro, Perplexity, and Perplexity Pro in responding to drug-related clinical questions, and examined the effect of an Enhanced Task Translation (ETT) technique on chatbot performance, after English technical terms were embedded with Thai-language prompts. An experimental comparative design was employed. Two sets of Thai-language multiple-choice questions (MCQs), each comprising 120 items based on the Thai Pharmacy Licensing Examination, were administered: one standard set and one ETT set. Cardiology-focused clinical case scenarios were additionally presented using a SOAP-note format. Performance was assessed using a rubric adapted from the American Society of Health-System Pharmacists (ASHP) guidelines across four domains: question classification, source citation, evidence application, and communication. All four chatbots surpassed the 60% passing threshold: 72.08–83.33% for the standard set and 73.75–80.83% for the ETT set. Perplexity Pro demonstrated the highest MCQ performance, while the inclusion of English technical terms did not consistently improve results across models. In the clinical case assessment, ChatGPT Pro achieved the highest rubric score, showing strong evidence use and clinical reasoning. These findings suggest that AI chatbot competency in drug-related queries is broadly comparable to that of pharmacists. The ETT technique did not produce notable performance differences between Thai-only and mixed-language prompts. AI chatbots may support pharmacists by generating initial structured responses to drug information requests. The proposed evaluation framework provides a potential model for assessing AI competency in non-English settings and underscores the importance of jointly evaluating knowledge accuracy and clinical judgment to ensure safe integration into drug information services.

Inthira Kanchanaphibool, Panyanat Aonpong, Thanaphat Dabngoen et al. · 0 citations
Review Open access Jul 2026

Enhancing patient care with AI agents: integrating advanced chatbot technologies for improved healthcare delivery.

The findings suggest that integrating advanced AI techniques enhances the accuracy, responsiveness, and contextual relevance of AI-driven medical agents, suggesting this scalable and reliable system presents a viable solution to reduce healthcare workload, enhance patient engagement, and democratize access to trusted medical information.

Yasmine Abu Adla, A. Hajj · 1 citation
Review Open access Aug 2026

A real‐world analysis of AI chatbot performance for medicines information enquiries

The provision of medicines information (MI) services requires interpretation and clinical judgement of complex scenarios by pharmacists. To date, few studies have assessed the performance of artificial intelligence (AI) chatbots to assist pharmacists providing MI advice. To evaluate the performance and risk associated with two AI chatbots (Microsoft Copilot and Google Gemini) to answer medicines‐related questions. A sample of 20 questions answered by the local MI service in November 2023 was entered in the two chatbot applications in January 2024 (round 1) and May 2024 (round 2). All questions were preceded with the prompt ‘I'm a pharmacist’. Chatbot responses were evaluated by comparing with a reference answer given by the MI service using a consensus process in the domains of content, patient management, risk of patient harm, and follow up review. Ethical approval was granted by the Canterbury District Health Board Research Office (Reference no: 20311) and the study conforms with the Declaration of Helsinki. For the 20 questions answered by both chatbots, few of the round 1 responses ( n = 4 for Copilot and n = 2 for Gemini) were considered complete and with adequate information to commence patient management with no risk of harm. Most were incomplete ( n = 13 for Copilot and n = 15 for Gemini) regarding content, but none were high risk of causing harm. In round 1, four responses from Copilot and eight from Gemini were flagged for follow up review. There was no significant difference in performance between chatbots in round 1 (p = 0.68) or between rounds 1 and 2 (Copilot p = 0.25 and Gemini p > 0.99). Our study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response.

Duncan Yorkston, Tracey Borrie, Paul K L Chin · 0 citations
Open access Aug 2026

AI Chatbots as a source of Ramadan medication management advice for patients with diabetes: a multilingual comparative evaluation

Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots—ChatGPT, Google Gemini, and Microsoft Copilot—when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic. Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation–Diabetes and Ramadan (IDF-DAR) Guidelines using a 0–2 accuracy scale, a 0–4 completeness scale, and a 0–3 safety scale. Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness ( p = 0.068) and a language effect on safety ( p = 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20, p = 0.009) and completeness (κ = 0.29, p = 0.001) and showed a very low κ in the safety scale (κ = 0.02, p = 0.81). AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.

S. Alomair, Maryam Alsuwayq, Walla Alabbad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.