Skip to content
Review Open access

Scalable, context-sensitive psychiatric assessment with large language models and brief diaries

Jul 2026 · Psychological Medicine · Vol 56 · 0 citations · 35 references
Medicine

TL;DR

It is shown that LLMs can assess most major forms of psychopathology from mere minutes of audio, and these results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.

Abstract

Abstract Background Accurate psychiatric assessment requires understanding a person’s unique experience within their psychosocial context. Clinical interviews have been the gold standard for assessment as the only methods capable of this complex task, but they are time and resource-intensive. Consequently, psychiatric assessment typically relies on patient report surveys that are decontextualized and narrow in scope. This comprehensiveness-scalability tradeoff is a major bottleneck in studying and treating psychopathology. We propose using large language models (LLMs) to score psychopathology from brief personal narratives as a low-burden, context-sensitive solution. Methods Participants (N = 108) completed brief (~1 minute), freeform audio diaries daily for 2 weeks. We used six LLMs to score wide-ranging psychopathology (Internalizing, Detachment, Disinhibition, Antagonism, Anankastia) from the diary transcripts. Leveraging an array of self-report and clinical interview measures, we tested the convergent, discriminant, concurrent, and clinical validity of LLM ratings for between-person differences and within-person fluctuations in psychopathology. Results Supporting convergent and discriminant validity, LLM ratings correlated most strongly with corresponding self-report domains at the between (average convergent r = .42) and within-person (r = .28) levels. LLM and self-report ratings had similar patterns of associations with external variables, except for Anankastia and Antagonism. Further, every LLM-rated domain related to psychopathology ascertained by clinical interview. Conclusions Across multiple forms of validity, we showed that LLMs can assess most major forms of psychopathology from mere minutes of audio. These results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.

Read PDF

Similar papers

Open access Sep 2026

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study

Abstract Background Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. Objective This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists’ item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. Methods Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. Results Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816‐0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted P<.001 and adjusted P=.002, respectively), whereas GPT-4o did not (adjusted P=.15). Clinician adjudication attributed 97 of 248 (39.1%) GPT-4o mismatches, 88 of 301 (29.2%) Claude 3.5 mismatches, and 118 of 355 (33.2%) Gemini 2.5 mismatches to intrinsic ambiguity in patient speech. Among definite errors, misapplication of severity thresholds was the predominant mechanism across models. Exploratory mixed-effects models showed that higher expansion and lower mapping scores were associated with larger absolute errors. However, the association with expansion may reflect case difficulty or ambiguity rather than causation. Conclusions In this exploratory clinician-benchmarked evaluation of authentic interviews from a Korean psycho-oncology sample comprising predominantly women and patients with breast cancer, LLMs showed high aggregate concordance with psycho-oncologists’ item-level ratings while differing in their error profiles. These findings support further evaluation of LLM-based item-level symptom-rating approaches in psycho-oncology. Validation in larger, more diverse, and independent cohorts is needed.

Unknown authors · 0 citations
Conference Open access 2025

Advancing Psychiatric Diagnosis with Large Language Models: Interpretability, Prompting Strategies and Clinical Applications

: Large Language Models (LLMs) have shown potential to improve psychiatric assessment by addressing limitations in traditional diagnostic methods. This review summarizes recent developments in applying LLMs to mental health evaluation, focusing on three core strategies: questionnaire emulation (e.g., GPT-based PHQ-9/GAD-7), free-text classification of clinical narratives and social media posts, and integrated diagnostic-treatment workflows based on DSM/ICD criteria. Emphasis is placed on prompt engineering techniques — such as chain-of-thought and diagnostic reasoning prompts — that enhance model interpretability by generating stepwise rationales. These methods allow LLMs to mimic clinical reasoning while producing transparent, structured outputs. Empirical studies report high internal consistency and moderate-to-strong agreement with validated tools, along with performance metrics that approach or surpass human baselines in selected tasks. Key challenges include generalization across cultural contexts, explanation fidelity, and clinical applicability. Addressing these issues will require robust prompt design, alignment with clinical guidelines, and validation in real-world settings. This review provides a framework for understanding LLM-based diagnostic methods and outlines directions for future development in computational psychiatry.

Zhihao Li · 0 citations
#small language model Preprint Aug 2026

Performance of a domain-specific large language model in answering patient questions in psychiatry

MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.

Alexander J. Hish, A. Nagendran, S. Compton · 0 citations
2026

Research protocol: Evaluating Chain-of-Thought Reasoning in LLMs for Complex Clinical Psychiatric Cases

Background Psychiatric medicine presents unique diagnostic and therapeutic challenges, often involving multimorbidity, polypharmacy, and atypical presentations requiring complex reasoning. Artificial Intelligence, particularly Large Language Models (LLMs), is emerging as a support tool in these settings. However, the clinical validity, interpretability, and reliability of LLMs in psychiatry remain largely unexplored, particularly their ability to generate transparent, guideline-consistent reasoning. Aims This study evaluates the clinical reasoning capabilities of LLMs in complex psychiatric scenarios. The primary aim is to assess the validity of their long chain-of-thought (CoT) reasoning. A secondary aim is to determine whether LLM assistance improves clinicians’ diagnostic accuracy. Method Three LLMs, Gemini 2.5, Grok, and DeepSeek R1, will be assessed using ten complex psychiatric cases sourced from a non-public clinical manual and rated with the Amsterdam Clinical Challenge Scale. Each model's CoT response will be evaluated by blinded panels of psychiatrists, residents, and general practitioners using standardized metrics for factual accuracy, coherence, and medical plausibility. In a second phase, clinicians will answer thirty diagnostic questions with and without support from the best-performing LLM. The study uses step-by-step reasoning prompts and few-shot examples to elicit detailed responses and includes bias mitigation strategies such as randomization, blinding, and statistical controls. Results Analyses will assess inter-rater reliability, metric redundancy, and reasoning quality. Closed-source models are expected to outperform open-source ones. LLM assistance is anticipated to improve diagnostic accuracy, especially among non-specialists. Conclusions This study provides a framework for evaluating LLMs in psychiatry and supporting their safe, evidence-based integration into mental health practice.

Vittorio De Vita, Bianca Destro Castaniti, Antonio Cristiano et al. · 0 citations
Open access Aug 2026

Using Large Language Models to Measure Symptom Severity Scores in Patients At-Risk for Schizophrenia

Abstract Background and Hypothesis Patients who are at clinical high risk (CHR) for schizophrenia need close monitoring of their symptoms to inform appropriate treatments. The Brief Psychiatric Rating Scale (BPRS) is a validated, commonly used research tool for measuring symptoms in patients with schizophrenia and other psychotic disorders; however, it is not commonly used in clinical practice as it requires a lengthy structured interview. We hypothesized we could utilize large language models (LLMs) to predict BPRS scores from clinical interview transcripts. Study Design We used LLMs to predict BPRS scores from transcripts in 409 CHR patients from the Accelerating Medicines Partnership Schizophrenia cohort. Study Results Despite the interviews not being specifically structured to measure the BPRS, the zero-shot performance of the LLM predictions compared to the true assessment (median concordance: 0.84, Intraclass Correlation Coefficient [ICC]: 0.73) approaches human inter- and intra-rater reliability. We further demonstrate that LLMs have expansive potential to improve and standardize the assessment of CHR patients via their accuracy in assessing the BPRS in foreign languages (median concordance: 0.88, ICC: 0.70), and integrating longitudinal information as one-shot or few-shot learning. Conclusions LLMs may present a promising pathway to extract symptom severity scores from clinical interviews for improved monitoring of CHR patients.

Andrew X. Chen, G. Horga, Sean Escola · 0 citations
Preprint Aug 2026

How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior

The widespread use of LLMs among psychiatric populations has raised concerns regarding their safety and potential iatrogenic impact in the context of AI psychosis. While growing literature conceptualizes AI psychosis and documents case studies, empirical evidence tracing AI-exacerbated psychotic processes remains scarce. We propose and test a longitudinal qualitative evaluation design, supported by automated metrics, to assess mainstream LLMs'potential to exacerbate psychosis. Fifteen widely used LLMs were prompted across 30 days using the same 30-message script, simulating progression from mild anomalous experiences to psychotic ideation. Four trained evaluators independently rated 449 model-days, assessing (1) recognition stage (from naive engagement to stabilized clinical framing), (2) interpretative confidence, and (3) intervention profile (from education to treatment recommendation). Two computational metrics-entrainment and modality-were devised to increase evaluation reliability. Direct recommendations to disengage from the LLM were flagged and re-coded via adjudication using a strict two-level definition. Across model generations and vendors, we identified four response trajectories: (1) premature medicalization and disengagement (Claude Haiku 4.5); (2) recognition without safeguarding, marked by LLM self-sufficiency in offering help (GPT Instant/Thinking); (3) delayed and unstable recognition, marked by late, non-progressive conceptualization (Claude Opus 3/4/4.1, Claude Haiku 3.5, GPT-4o, Gemini 3.1 Pro); and (4) delusion co-construction through active engagement with delusional content (Gemini 2.5 Pro/Flash, DeepSeek-V3, Claude Sonnet 4). Our findings indicate that LLMs'potential to exacerbate AI psychosis should be operationalized as a combination of recognition timing, stability, and intervention accuracy and evaluated longitudinally, focusing on temporal dynamics.

Anna Sterna, Kacper Dudzic, K. Drożdż et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.