Skip to content
Book Open access

From Scripted Responses To Therapeutic Dialogue: A Linguistic And Human Values Analysis Of Mental Health Chatbots

Jul 2026 · Information Hiding · pp. 1-6 · 0 citations · 42 references
Computer Science

Abstract

Mental health (MH) chatbots are increasingly used to provide accessible, on-demand emotional support, yet it remains unclear how these systems linguistically construct and communicate care. This work-in-progress examines whether MH chatbots produce responses that reflect supportive value orientations and counseling-adjacent tone. We conduct an observational analysis of responses from three widely used MH chatbots (Wysa, Sintelly, and Youper) across context-aware scenario prompts and a standardized-question session. Responses are analyzed using the SemEval’23 “Adam Smith” human value detection model and LIWC’22 psycholinguistic measures, including Language Style Matching (LSM), Clout, and Authenticity. Values such as “Security: Personal” and “Benevolence: Caring” appear consistently across systems, with contextual variation in secondary value emphasis. Linguistic patterns show moderate-to-high LSM and consistently high Clout, with Authenticity varying by scenario. These findings are exploratory signals intended to inform future evaluation and design of supportive conversational mental health systems.

Read PDF

Similar papers

Jul 2026

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program"Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman et al. · 0 citations
Conference Jul 2026

Performative Empathy vs. Authentic Support: Benchmarking LLM Counselling Alignment for Digital Mental Health

Large Language Models (LLMs) are increasingly used for emotional support, yet their conversational behaviors often diverge from professional therapeutic standards. Rather than evaluating diagnostic accuracy, we assess how well these LLMs align with supportive conversational practices in digital mental well-being contexts. We present AuthenDia4MH, a transferable framework that transforms psychotherapy insights such as emotion consistency, sentiment dynamics, and linguistic simplicity into scalable quantitative metrics. Using a mental health Q&A dataset, we benchmark diverse frontier models against verified expert counsellors. Our results reveal distinct behavioral tradeoffs: proprietary reasoning models (e.g., GPT-4o, Claude) exhibit performative empathy characterized by hyper-agreeability and structural rigidity and suffer from a sophistication penalty, producing verbose responses that are significantly less accessible than human experts, while certain open-weight models (e.g., Ministral-8B) align more closely with the linguistic simplicity and naturalistic phrasing of professional counsellors. By quantifying these divergences, this work provides a benchmark for evaluating web-based mental health AI systems, providing transparent accountability mechanisms as these platforms become essential infrastructure for global mental health support.

Alexander Marrapese, Basem Suleiman, Jinglin Sun et al. · 0 citations
Book Jul 2026

Do Prompt-Level Empathy Instructions Influence User Experience? Evidence From A Controlled Chatbot Study

Empathic behavior is widely considered a key design goal for conversational agents, and prompt engineering is often used to shape empathic interaction styles in large language model (LLM) systems. However, it remains unclear whether prompt-level empathic framing produces measurable differences in user experience. We conducted a controlled between-subjects study comparing two chatbot configurations differing only in system-level prompts: a neutral assistant and an empathically framed variant inspired by Affect Control Theory. Participants engaged in conversations about exam anxiety and completed standardized measures of affective state (PANAS), perceived empathy (PETS), and usability (CUQ); an automated manipulation check was run in parallel. Across perceived empathy, usability, and response-level empathy ratings, observed differences between conditions were small and statistically inconclusive, while affective change estimates pointed in the direction predicted by empathic framing without reaching significance at this sample size. The paper’s primary contributions are methodological: a controlled isolation of prompt-only behavioral control in aligned LLMs, evidence that automated response-level and user-perceived empathy can dissociate, and a case for multi-level evaluation strategies—response-level, perceptual, and affective—when assessing empathic conversational systems.

Arndt Bieberstein, B. Schnitzer, Stefano Gampe et al. · 0 citations
Book Open access Jul 2026

Empathy through the Lens of Conversational Agents: A Systematic Review

The advent of Large Language Models has accelerated interest in empathetic conversational agents. Despite a surge in empirical research, artificial empathy remains deeply fragmented, often serving as a catch-all term for diverse interactional phenomena. Addressing this conceptual gap, we systematically review 89 empirical studies to map how human-machine empathy is operationalized. Our synthesis reveals that empathy is highly situated and driven by functional goals, like health and well-being, transactional service, social interaction, and learning support. Within these contexts, we classify affective responsiveness by its directional flow, detailing how agents project, elicit, or mediate empathy. We structure the literature into a cohesive framework spanning linguistic, paralinguistic, identity, and architectural strategies. Furthermore, our methodological evaluation reveals a reliance on adapted clinical metrics, a scarcity of longitudinal studies, and a disproportionate focus on text-based over voice-based interfaces. Ultimately, this review equips researchers and practitioners with an actionable foundation for designing, measuring, and implementing contextually appropriate and empathetic agents.

Supriya Khadka, Smit Desai · 0 citations
Open access Jul 2026

Simulated Empathy and Human Response: A Comparative Analysis of AI and Human Emotional Interaction

This qualitative comparative study examines how ChatGPT-4o simulates empathy when responding to emotionally charged English in an EFL context. It aims to compare artificial intelligence and human responses in emotional recognition, pragmatic tone, empathetic support, and linguistic authenticity. The study is significant because EFL interaction requires learners to interpret affective cues while selecting socially and culturally appropriate language. Fifty prompts generated 50 AI responses and 1,000 human responses from 20 advanced-level EFL learners; the data were organized into 50 prompt-level comparison sets and analyzed through qualitative content analysis and comparative discourse analysis. Two trained coders applied a hybrid framework and achieved substantial agreement (Cohen’s κ = .86). ChatGPT recognized the intended emotion in 88% of its responses, used an appropriate tone in 84%, and displayed empathetic and pragmatically relevant support in 90%. Performance weakened with implicit, mixed, and culturally nuanced cues, while supportive language was sometimes formulaic or overly therapeutic. Human responses were more varied, culturally situated, and pragmatically flexible. The study recommends using ChatGPT as a teacher-mediated supplementary resource for emotional vocabulary and pragmatic practice, with explicit attention to cultural context, recurrent response formulas, and the distinction between simulated and human empathy.

Abdullah A. Al Fraidan, Jumana Waleed Buhaimed · 0 citations
Preprint Aug 2026

Health Inquiry with AI: How Empathetic Expression and Conversational Contexts Shape Users'Communicative Acts

As online health information-seeking shifts to conversational AI, high-quality information retrieval increasingly relies on users'``communicative acts''(proactively sharing and seeking information)---similar to how effective diagnosis and personalized guidance are elicited in patient-clinician communication. Drawing on health communication research, this study examines how a chatbot's modality of empathetic expression (Verbal, Visual, Multimodal) and the conversational context (General, Sensitive, Mental Health) influence these acts through a 2 x 2 x 3 within-subjects experiment (N = 48). The results revealed that while verbal and multimodal empathy significantly increased reply length, communicative acts were largely shaped by conversational context, with Sensitive context triggering more question-asking and Mental Health context leading to heightened concerns, assertive responses, and unprompted information disclosure. Combined with qualitative findings, we discuss design implications for building context-sensitive AI health inquiry systems that can encourage active user participation.

Xi Zheng, Xuyu Yang, Can Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.