Psychiatry’s reliance on language makes LLMs a natural tool for psychopathological assessment, yet structured, item-level assessments from psychiatric clinical interviews remain under-researched. In this proof-of-concept study, 10 LLMs assessed transcripts of three simulated psychiatric interviews across all 100 items of the Association for Methodology and Documentation in Psychiatry (AMDP) system, benchmarked against 108 early-career clinicians rating full audiovisual recordings, using an expert consensus panel as reference. GPT-5.1 and Gemini-3-Pro-Preview achieved the highest accuracy (0.72; 64th percentile of the clinician distribution) using majority voting across three runs with AMDP definitions as context. GPT-5.1, selected for a marginal advantage, showed per-scenario accuracies of 0.81 (depression), 0.76 (mania), and 0.60 (schizophrenia) versus clinician means of 0.79, 0.68, and 0.58. Clinicians and LLMs showed distinct error profiles: clinicians tended to over-infer symptom presence, whereas LLMs more conservatively flagged items as “not assessable” — most pronounced for observation-dependent items but present even for text-assessable items (19.4% vs. 11.4%, p < 0.001). In post hoc simulated disagreement resolutions (2091 clinician pairs; 35.5% disagreements), LLM and board-certified supervision were associated with more accurate resolutions than unsupervised random clinician selection (p < 0.0002). These proof-of-concept findings require validation in real patient interviews, larger samples, and prospective studies integrating multimodal input.
: Large Language Models (LLMs) have shown potential to improve psychiatric assessment by addressing limitations in traditional diagnostic methods. This review summarizes recent developments in applying LLMs to mental health evaluation, focusing on three core strategies: questionnaire emulation (e.g., GPT-based PHQ-9/GAD-7), free-text classification of clinical narratives and social media posts, and integrated diagnostic-treatment workflows based on DSM/ICD criteria. Emphasis is placed on prompt engineering techniques — such as chain-of-thought and diagnostic reasoning prompts — that enhance model interpretability by generating stepwise rationales. These methods allow LLMs to mimic clinical reasoning while producing transparent, structured outputs. Empirical studies report high internal consistency and moderate-to-strong agreement with validated tools, along with performance metrics that approach or surpass human baselines in selected tasks. Key challenges include generalization across cultural contexts, explanation fidelity, and clinical applicability. Addressing these issues will require robust prompt design, alignment with clinical guidelines, and validation in real-world settings. This review provides a framework for understanding LLM-based diagnostic methods and outlines directions for future development in computational psychiatry.
Zhihao Li· Proceedings of the 3rd Inter...· 0 citations
It is shown that LLMs can assess most major forms of psychopathology from mere minutes of audio, and these results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.
Whitney R. Ringwald, Aman Taxali, Michael Angstadt et al.· Psychological Medicine· 0 citations
MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Alexander J. Hish, A. Nagendran, S. Compton· 0 citations
Background
Psychiatric medicine presents unique diagnostic and therapeutic challenges, often involving multimorbidity, polypharmacy, and atypical presentations requiring complex reasoning. Artificial Intelligence, particularly Large Language Models (LLMs), is emerging as a support tool in these settings. However, the clinical validity, interpretability, and reliability of LLMs in psychiatry remain largely unexplored, particularly their ability to generate transparent, guideline-consistent reasoning.
Aims
This study evaluates the clinical reasoning capabilities of LLMs in complex psychiatric scenarios. The primary aim is to assess the validity of their long chain-of-thought (CoT) reasoning. A secondary aim is to determine whether LLM assistance improves clinicians’ diagnostic accuracy.
Method
Three LLMs, Gemini 2.5, Grok, and DeepSeek R1, will be assessed using ten complex psychiatric cases sourced from a non-public clinical manual and rated with the Amsterdam Clinical Challenge Scale. Each model's CoT response will be evaluated by blinded panels of psychiatrists, residents, and general practitioners using standardized metrics for factual accuracy, coherence, and medical plausibility. In a second phase, clinicians will answer thirty diagnostic questions with and without support from the best-performing LLM. The study uses step-by-step reasoning prompts and few-shot examples to elicit detailed responses and includes bias mitigation strategies such as randomization, blinding, and statistical controls.
Results
Analyses will assess inter-rater reliability, metric redundancy, and reasoning quality. Closed-source models are expected to outperform open-source ones. LLM assistance is anticipated to improve diagnostic accuracy, especially among non-specialists.
Conclusions
This study provides a framework for evaluating LLMs in psychiatry and supporting their safe, evidence-based integration into mental health practice.
Vittorio De Vita, Bianca Destro Castaniti, Antonio Cristiano et al.· Monolith alpha· 0 citations
Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients'functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts (U = 241, p<.001, r = .62; LLMs: M = 41.43, SD = 1.99; experts: M = 29.03, SD = 8.05), but it was unrelated to decision consistency or functional assessment. While LLMs and physicians granted support less frequently than psychologists (U = 180.50, p = .003, r = .34), they also interpreted a concept of"basic life needs"differently, primarily as biological survival and self-care, and not communicative and social needs. These findings suggest that despite expressing high levels of intellectual humility, LLMs reproduce a reductionist interpretive framework and knowledge embedded in medical decision-making. More broadly, we argue that evaluating AI in high-stakes contexts requires not only measuring accuracy, agreement, or resistance to cognitive bias, but also critical examination of the concepts of neurodiversity that AI systems operationalize.
Maciej Wodziński, Joanna Wodzińska, Kacper Dudzic et al.· 0 citations