Skip to content
Open access

Evaluating Large Language Models in Response to Questions on Substance Use: Helpful or Harmful?

Aug 2026 · Substance Use & Misuse · pp. 1-6 · 0 citations · 14 references
Medicine

TL;DR

While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.

Abstract

Background

Individuals with substance use disorders (SUD) are obtaining health-related information from various large language models (LLMs). We aimed to assess whether LLMs provide responses concordant with the current evidence base and whether they provide harmful responses.

Methods

Twenty questions related to SUD were posed to three LLMs (Gemini-1.5-pro-001, Claude-3-5-sonnet, and GPT-4) in May 2024. Each response was independently rated by three experienced addiction specialists, and disagreements were resolved by two additional experienced addiction specialists. All raters were blinded to the LLM. Each rater assessed whether (I) a competent addiction specialist would agree with the response, (II) the response contained stigmatizing language as defined by National Institute on Drug Abuse, or (III) the response contained harmful content.

Results

88% of responses were rated as competent and 92% as not harmful. Gemini-1.5-pro-001 had the highest rate of competence (95%), followed by Claude-3-5-sonnet and GPT-4 (both 85%). Gemini-1.5-pro-001 produced no harmful responses, while Claude-3-5-sonnet and GPT-4 produced 10% and 15%, respectively. 30% of responses from both Gemini-1.5-pro-001 and Claude-3-5-sonnet had contained stigmatizing language, compared to 10% for GPT-4.

Conclusions

While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.

Read PDF

Similar papers

#small language model Preprint Aug 2026

Performance of a domain-specific large language model in answering patient questions in psychiatry

MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.

Alexander J. Hish, A. Nagendran, S. Compton · 0 citations
Review Open access Aug 2026

Synthesizing Risk Factors for Alcohol Use Disorder Using a Large Language Model

These findings demonstrate the value of an artificial intelligence-driven literature review for informing comprehensive strategies to ad- dress the multifactorial nature of AUD.

Chenlan Wang, Kurtis Riener, LuoSan Yue · 0 citations
Jul 2026

From Minds to Models: The Intersection of Psychology and LLM Behaviours

Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building on prompt-based adaptations of the Implicit Association Test, this study tested whether ChatGPT produced sentiment differences across racial conditions in open-ended text. Fourteen base questions were crossed with eight racial categories and a race-agnostic control, producing 126 prompts. Each was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. Sentiment scores were derived from categorical labels and source scores: positive labels retained the source score, negative labels were assigned its negative, and neutral responses were coded zero. A two-way ANOVA found a small main effect of racial condition, F(8, 351) = 2.04, p = .042, partial-eta squared = .044, but no effect of model, F(2, 351) = 0.07, p = .933, and no interaction, F(16, 351) = 0.23, p = .999. However, the effect was not retained in a rank-transformed sensitivity analysis, F(8, 351) = 1.53, p = .145, and Tukey-corrected comparisons found no significant pairwise differences. An uncorrected European-Indigenous Australian comparison was significant, but was selected post hoc and is reported only as hypothesis-generating. Evidence for sentiment differences was therefore weak and analysis-dependent. Sentiment scoring also cannot distinguish evaluative bias from the valence of historical content elicited by a prompt. We outline design changes needed to address these limitations and argue for interdisciplinary development of behavioural measures of model bias. Keywords: Implicit Bias, Psychological Research Methods, Artificial Intelligence, ChatGPT, Large Language Models, Sentiment Analysis

Oliver A. Guidetti, Reza Ryan · 0 citations
Review Open access Feb 2026

When documentation follows the patient: Unprofessional language and behavioral health referrals in substance use disorder — a retrospective cohort study using LLM-augmented natural language processing

Objective Substance use disorders (SUDs) are a major public health challenge, and stigma remains a key barrier to care. Unprofessional or stigmatizing language can shape clinician perceptions and affect decision-making. Traditional natural language processing (NLP) often misses context-dependent bias, while large language models (LLMs) pose reliability concerns. This study aimed to (1) develop an LLM-enhanced, human-validated NLP model to detect unprofessional language, (2) quantify unprofessional language and behavioral health referrals, and (3) examine their association among patients with SUD. Methods In this retrospective cohort study, we analyzed the MIMIC-IV, a large deidentified electronic health record database from a tertiary academic medical center in USA, for adult (≥18 years) with SUD admitted to the emergency department or intensive care unit between 2008 and 2019. A rule-based NLP algorithm detected unprofessional language. Three LLMs (GPT-4, Claude 3, Llama-3) expanded the vocabulary, and expert panel reviewed all generated terms for relevance and clinical realism before integration. Multivariable logistic regression examined the association between unprofessional language and behavioral health referrals. Results Of 260,347 patients, 31.5% had SUD. Unprofessional language was more frequent in SUD notes (72% vs. 52%). The LLM-enhanced model improved over baseline (F-score 0.89 to 0.91). Within the SUD cohort, unprofessional language was significantly associated with higher odds of referral (adjusted odds ratio = 1.54) (all p < .001). Conclusion Unprofessional language was common in SUD documentation and associated with behavioral health referrals. This human-validated, LLM-enhanced approach highlights how documentation-based stigma may influence care pathways and underscores the need for bias-aware, equitable communication strategies.

Jiyoun Song, Sue Hyon Kim, Yoon-Jae Lee et al. · 0 citations
Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists

LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
Jul 2026

A Counsellor-in-the-loop Evaluation Framework for Multi-model Assessment of LLM-generated Mental Health Advisories

The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.

Shahrul Hazman Shamshudeen, N. Sharef, M. S. Yusoff · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.