Skip to content
Review Open access

Real-world use of large language models for mental health in 2024

Aug 2026 · npj Digital Medicine · Vol 9 · 1 citation · 67 references
Medicine

Abstract

The extent to which people use general-purpose large language models (LLMs) for their mental health is unknown. Information about use patterns is important for clinicians, developers, and regulators. We surveyed U.S. adults (n = 1871) between August and October 2024 using stratified sampling across age, sex, and race/ethnicity to approximate national demographics. We found that 24% of participants use LLMs for mental health; they are disproportionately young, male, and Black, and have poor mental health. Participants reported difficulty accessing traditional treatment and using LLMs because they are free, convenient, and available. They report using LLMs for emotional support, learning therapy skills, and supplementing existing therapy. Using Pew-reported estimates of population LLM use, we conservatively estimate that as of 2024, 14–18 million U.S. adults may have been using LLMs for mental health. This work highlights the need for monitoring and evaluation to understand the potential harms and benefits of such use.

Read PDF

Similar papers

Review Aug 2026

Large language model applications for real-time clinical mental health assessment: Current potential and future directions.

Large language models are best understood as emerging assessment-support tools rather than replacements for clinical evaluation because the limited pace of academic validation means that, at present, LLMs are best understood as emerging assessment-support tools rather than replacements for clinical evaluation.

K. Aafjes-van Doorn, Francine Cheng Ty, A. Hua et al. · 0 citations
Open access Aug 2026

The use of large language models in automated depression detection.

BACKGROUND Large language models have been evaluated on many healthcare tasks, including depression screening. However, it is unclear whether estimates of performance are accurate, especially in a setting with realistic clinical constraints. METHODS We use the publicly-available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) to give performance estimates of LLMs that respect patient privacy and could be feasibly deployed in a clinical setting. These models are locally run and under 15 billion parameters. RESULTS Accuracy, sensitivity, and specificity of the models we evaluated ranged from 0.233-0.677, 0.041-0.929, 0.000-0.729, respectively. There are significant differences in the performance we observed versus other studies that evaluate commercial models. We also demonstrate poor agreement amongst different LLMs. CONCLUSION Current performance estimates of LLMs with respect to depression screening are most likely optimistic. When restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.

S. H. Ling, W. Chorney · 0 citations
Review Open access Aug 2026

Generative AI Use for Mental Health Support: Patterns, Correlates, and Impact among Canadian Students

Background. General purpose generative AI (GenAI) chatbots are increasingly used by students for mental health support. Research on prevalence estimates vary widely, rarely link use to validated clinical measures, and have not been reported in a Canadian student population. We estimated the prevalence trends, patterns, perceived impact, and correlates of GenAI use for mental health support among Canadian university students. Methods. We analysed one year (May 2025 to April 2026) repeated cross-sectional data from the Canadian arm of the WHO World Mental Health International College Student survey (WMH-ICS) The primary outcome was past-year prevalence of GenAI use for mental health support. Specific use purposes, perceived impact, reasons for non-use, and future use intent were also analysed. Factors associated with GenAI use were assessed using modified Poisson regression. As a sensitivity analysis, an elastic-net penalised regression model was fitted to assess the robustness of findings to an alternative modelling approach. Results. The past-year prevalence of GenAI chatbot use for mental health support was 25.2% (95% CI: 22.7 - 27.9), with a lifetime prevalence of 30.2%. Use was mostly occasional and predominately for seeking mental health information, stress management, and emotional support/companionship. Students of Asian ethnicity, those with higher clinical burden, recent adverse life experiences, weaker social support, and prior digital help-seeking behaviours were more likely to use GenAI for mental health purposes. Conversely, 2SLGBTQ+ students and those with romantic partners were less likely. Nearly three-quarters (74.2%) of users perceived such use to have a positive impact on their mental health and emotional wellbeing. Non-users reported preference for human interaction, distrust of GenAI in mental health (67.4% each), and privacy/security concerns (50.3%). Non-use also reflected principled objections to AI, including ethical and environmental concerns, with most non-users indicating no future use intention. Conclusions. GenAI chatbot use for mental health support has become commonplace among Canadian university students and is concentrated among those with greater mental health needs and fewer social support resources. Although most users perceived these tools as beneficial, their clinical effectiveness and safety remain uncertain. Rigorous prospective studies are needed to determine whether perceived benefits translate into improved mental health outcomes and whether purpose-built GenAI mental health interventions offer greater clinical benefit and safety than general-purpose chatbots.

L. Olisaeloka, R. Munthali, D. Vigo · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Book Open access Jul 2026

Large Language Models and the Evolution of Online Help-Seeking for Mental Health

Large language models (LLMs) are rapidly becoming embedded in everyday mental health help-seeking practices, particularly among young people who already turn to digital platforms as gateways to mental health support. While LLMs offer unprecedented immediacy and accessibility, their integration into help-seeking ecosystems raises important questions for digital health research. This opinion paper argues that LLMs fundamentally reshape the developmental processes underpinning online help-seeking. Traditional digital help-seeking requires active exploration, searching, comparing sources and reflecting on lived experience, processes that contribute to mental health literacy and resilience. In contrast, LLMs collapse informational plurality into singular, authoritative-sounding responses, potentially shifting users from active exploration toward passive consumption. We discuss the risks of sycophancy, and over-reliance on immediacy, and consider how these dynamics may alter developmental trajectories of coping and help-seeking agency. We argue that preserving agency, connectedness, and reflective engagement must be central to the design of conversational AI in health contexts.

Claudette Pretorius · 0 citations
Editorial Open access Jul 2026

Digital Mental Health Research Priorities, Revisited for the AI and Large Language Model Era

Abstract Digital mental health has become an established part of mental health care, but the rapid arrival of large language models and other artificial intelligence (AI) tools has refocused attention on the evidence needed to guide the field. This editorial updates the research priorities articulated by JMIR Mental Health in 2023, while reaffirming their emphasis on equity, replicability, privacy, efficacy, and engagement. While the importance of these priorities has not changed in recent years, the urgency with which they must now be applied has. As digital tools become more clinically consequential, research must move beyond demonstrating that a technology is feasible, usable, or novel. The field now needs studies that clarify how these tools work, for whom they are beneficial, under what conditions they may cause harm, and how they can be ethically integrated into care. We call for research that is transparent about the technologies being studied, grounded in meaningful clinical questions, attentive to safety, and designed to produce knowledge that remains useful as specific products and models change.

M. Birk, Shruti Kochhar, Keris Myrick et al. · 1 citation