Skip to content
Open access

607. When artificial intelligence becomes a rater: detecting formal thought disorders in schizophrenia

Sep 2026 · International Journal of Neuropsychopharmacology · Vol 29, pp. i56 - i56 · 0 citations

Abstract

Abstract Background Formal thought disorders (FTDs) are a fundamental component of the psychopathology of schizophrenia (SCZ), characterized by fragmented conceptual organization and loss of associative coherence [1]. FTDs are more frequent and severe in individuals with treatment-resistant schizophrenia (TRS), who show poorer functional outcomes and higher disease burden [2]. Accurate evaluation of FTDs is essential for monitoring symptom severity and assessing treatment or rehabilitation effects. The Thought and Language Disorder (TALD) [3] scale supports a structured assessment of FTDs but remains time-consuming and influenced by clinician judgment. With advances in natural language processing (NLP), large language models (LLMs) offer the opportunity to support an objective, reproducible, and scalable assessment of language disturbances in psychiatry. Aims & Objectives This study aims to determine whether an LLM can reproduce clinician scoring of the TALD in individuals with SCZ and TRS, and to assess the agreement between automated and expert ratings. To our knowledge, this is the first study to perform a full-scale, item-level TALD scoring through an LLM. Method Thirty-three patients (14 TRS, 19 SCZ) underwent a recorded clinical interview, which was independently rated by two psychiatrists and an LLM trained on TALD scoring. TALD items were converted into quantitative metrics through two analytic frameworks: a time-stamp analysis (TSA) module to quantify temporal speech features (for example, pressured speech was assessed as pauses and words per minute), and a NLP module to assess semantic parameters (e.g., perseveration was converted into the number of repeated responses out of the total number of responses given by the patient). A mixed-design repeated-measures ANOVA tested effects of Rater (clinician vs LLM) and Group (SCZ vs TRS). Agreement was evaluated via intraclass correlation coefficient (ICC, two-way mixed, absolute agreement) for total scores and weighted Cohen’s κ for item-level concordance. Results The LLM consistently assigned slightly higher total TALD scores than clinicians (clinicians: 26.4 ± 10.7; LLM: 28.0 ± 8.0). ANOVA revealed a significant main effect of Rater (p = 0.001), with no significant Group effect and no Rater × Group interaction. Agreement for total TALD scores was good (ICC = 0.84), comparable across SCZ (ICC = 0.86) and TRS (ICC = 0.83) (Table 1). Weighted κ demonstrated moderate-to-almost-perfect concordance for most items (e.g., blockage κ ≈ 0.94), while agreement was lower for infrequent phenomena such as dissociation of thinking (κ ≈ 0.33), likely reflecting limited occurrence and broader variability in clinician scoring. Discussion & Conclusions LLM assigns slightly higher scores than clinicians - potentially reflecting the model’s lack of emotional or cultural biases. Overall agreement is good, supporting its potential role in standardising psychopathology evaluation. However, such tools are to be used under medical supervision and require further clinical validation before integration into clinical practice. With appropriate safeguards, LLM-assisted scoring could enhance reproducibility in psychiatric research and offer scalable support in the clinical assessment of SCZ and TRS.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.