Beyond Labels: Training Cognitive Empathy with Context-Rich Synthetic Dialogues
Abstract
As Large Language Models (LLMs) are increasingly deployed in emotionally sensitive applications like mental health support, the need for genuine artificial empathy has become critical. Existing conversational datasets often fail to instill deep cognitive empathy, instead promoting superficial affective mimicry. To address this gap, we introduce RelationalDialogues, a novel, fully synthetic dataset of 12,849 multi-turn dialogues designed to explicitly train perspective-taking. Using a 120-billion parameter LLM, we developed a data generation pipeline that grounds each conversation in structured metadata, including distinct speaker/listener backgrounds, situational stimuli, and relationship dynamics derived from a psychologically justified emotion taxonomy. A key feature of our methodology is blinding the "listener" agent to the explicit emotion label during generation, forcing it to infer emotional states from conversational context alone. We fine-tuned a Llama-3-8B-Instruct model on this corpus and evaluated it against its base counterpart using an LLM-as-a-Judge framework across three dimensions: Emotion Recognition, Perspective-Taking, and Emotional Contagion. Our fine-tuned model achieved a 51.05% win rate over the base model (45.53%) and demonstrated statistically significant improvements, particularly in Perspective-Taking (average score increase from 3.06 to 3.22) and Emotional Contagion (2.96 to 3.09). Granular analysis reveals the fine-tuned model excels at navigating nuanced, socially complex emotions like loneliness and jealousy, whereas the base model relies on pre-programmed templates for high-arousal emotions. These findings demonstrate that training on highly contextualized, metadata-driven synthetic data is an effective method for advancing LLMs from displaying superficial sympathy to engaging in genuine cognitive perspective-taking, a critical step for developing more emotionally intelligent conversational agents.