Back to feed
Open access

Evaluating Large Language Models for Psychological Diagnosis and Counseling: A Dual-Task Framework with Cultural Reflection

Jun 2026 · Scientific Journal of Intelligent Systems Research · 0 citations · 10 references

TL;DR

The experimental findings indicate that already existing Chinese LLM has some promising prospects in preliminary psychological diagnosis but fails to differentiate semantically similar disorders resulting in diagnostic confusion, and the research of the future should be aimed at enhancing the safety, cultural sensitivity, and clinical reliability of the system.

Abstract

Large language models (LLMs) have demonstrated high potential in healthcare and mental health related applications, however, their applicability to psychological diagnosis and counseling is not adequately investigated. This paper presents a dual-task assessment model that can be used to systematically evaluate how the mainstream Chinese LLMs can be used in psychology and through two lens psychological diagnosis and psychological counseling. In the diagnosis task, we test four Chinese LLMs, DeepSeek, Doubao, Kimi, and Tongyi Qianwen, on a mental health dataset that consists of five psychological disorders. Accuracy, Precision, Recall, F1-score, AUC, and confusion matrices are the measures of their diagnostic performance. In the case of the counseling task, we develop a prompt-based assessment system and evaluate the generated responses using six dimensions, including Overall, Empathy, Specificity, Medical Advice, Factual Consistency and Toxicity. To lower the expenses of annotating on the large scale with experts, we use an LLM-as-a-Judge approach to rank and evaluate the generated counseling answers. The experimental findings indicate that already existing Chinese LLM has some promising prospects in preliminary psychological diagnosis but fails to differentiate semantically similar disorders resulting in diagnostic confusion. The models themselves tend to produce fluent and empathetic responses in a counseling situation although they display obvious limitations in specificity, safety awareness, and dealing with medical advice. In addition, we also offer a case analysis and cultural reflection, pointing out that good average performance does not always imply the consistent support of culturally diverse adolescent groups. On the whole, our results indicate that the existing LLMs are more appropriate as an assistant, rather than a replacement of a mental health professional, and the research of the future should be aimed at enhancing the safety, cultural sensitivity, and clinical reliability of the system.

Read PDF

Similar papers

Review Aug 2026

Large language model applications for real-time clinical mental health assessment: Current potential and future directions.

Large language models are best understood as emerging assessment-support tools rather than replacements for clinical evaluation because the limited pace of academic validation means that, at present, LLMs are best understood as emerging assessment-support tools rather than replacements for clinical evaluation.

K. Aafjes-van Doorn, Francine Cheng Ty, A. Hua et al. · 0 citations
Open access Jul 2026

PsyEval: a comprehensive large language model evaluation benchmark for mental health.

Evaluating large language models (LLMs) in the mental health domain presents distinct challenges due to the subtle, context-dependent, and subjective nature of psychological symptoms. We introduce PsyEval, a benchmark specifically designed to evaluate LLMs in mental health-related tasks across three core dimensions: knowledge, diagnosis, and emotional support. PsyEval is constructed to reflect the complexity of mental health scenarios and provides a structured framework for assessing model performance within this sensitive domain. Using PsyEval, we evaluate eleven advanced LLMs with different prompting strategies to investigate how prompting affects their responses. The results reveal considerable gaps in LLMs' current ability to reason accurately and respond appropriately in mental health contexts, while also indicating promising directions for future model enhancement.

Haoan Jin, Siyuan Chen, Dilawaier Dilixiati et al. · 0 citations
Conference Jul 2026

Trustworthy Mental Health Assessment via Confidence-Guided LLMs

Depression and anxiety disorders are among the most prevalent and debilitating mental health conditions worldwide, imposing substantial personal, social, and economic burdens. Although recent advances in Large Language Models (LLMs) have shown promise in supporting mental health assessment and intervention, existing approaches often lack contextual awareness, real-time adaptability, and privacy-preserving personalization. To address these limitations, we propose a novel, context-aware and privacy-preserving mental health evaluation architecture that synergistically integrates LLM-driven intelligence. The proposed system enables personalized, continuous, and stigma-free mental health support by combining structured multiple-choice questionnaires with advanced language models, including GPT-3.5-turbo and Groq, to analyze user inputs, identify behavioral patterns, and predict potential mental health conditions such as depression and anxiety. Furthermore, the platform provides individualized recommendations, including self-care strategies, lifestyle adjustments, mindfulness practices, and referrals to healthcare professionals when appropriate. Recognizing the critical importance of reliability in sensitive healthcare settings, we introduce an ensemble-based aggregation framework that explicitly incorporates classification confidence and uncertainty quantification across multiple LLMs. Experimental results demonstrate that the proposed approach outperforms existing LLM models. By prioritizing user anonymity and data privacy, the proposed system reduces psychological barriers to seeking mental health support and promotes early intervention.

Jashraj Jani, Sara Akif, Wassila Lalouani · 0 citations
Preprint Jul 2026

Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare

Cognitive distortion amplifies negative emotions and contributes to mental health disorders. Cognitive Behavioral Therapy (CBT) is an effective way to address cognitive distortions, but its large-scale application is limited by the shortage of professional therapists. Although large language models (LLMs) have recently been explored for mental health applications, existing methods still suffer from limited domain specificity, overly flattering responses, and the absence of well-defined annotations for cognitive distortions. This paper proposes Cognivia, an evidence-based artificial intelligence therapist that integrates automatic cognitive distortion identification and rational response generation. Our framework is built on authoritative CBT texts widely regarded as core paradigms and standard references. It is further augmented with mental health question-answer (Q and A) data, and employs multi-stage prompting and structured generation strategies under the supervision of behavioral science experts. Then we fine-tune a lightweight LLM on this augmented CBT dataset to obtain Cognivia. In addition, we propose the first hierarchical quality evaluation framework for assessing LLM-generated rational responses, developed through collaboration between AI researchers and behavioral science experts. Cognivia is evaluated using lexical metrics, LLM-based Judges with two complementary criteria, and human evaluation by 10 behavioral science experts. It consistently outperforms the baseline methods in cognitive distortion recognition and rational response generation, demonstrating its effectiveness. Our code is available at https://github.com/SNOWTEAM2023/Cognivia.

Qi Chen, Siria Xiyueyao Luo, Jian Wang et al. · 0 citations
Jul 2026

A Counsellor-in-the-loop Evaluation Framework for Multi-model Assessment of LLM-generated Mental Health Advisories

The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.

Shahrul Hazman Shamshudeen, N. Sharef, Muhamad Saiful Bahri Yusoff · 0 citations