Skip to content

From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue

Jul 2026 · arXiv.org · Vol abs/2607.07824 · 0 citations · 37 references
Computer Science

TL;DR

CPM-MultiAgent is proposed, a CPM-grounded emotion evolution multi-agent framework for supporting emotional changes in persona-based dialogue that represents a character's emotion as a latent state that is continuously reshaped by dialogue triggers.

Abstract

Large Language Models (LLMs) have substantially advanced persona-based dialogue agents for emotion-sensitive role simulation in healthcare, education, counseling, customer service, and interactive storytelling. However, two related lines of work leave a key gap. Persona-based dialogue systems often encode emotions as static traits or surface-level stylistic cues, and affective dialogue research has largely focused on empathetic response generation toward users rather than modeling the agent persona's own evolving emotional state. As a result, trigger-driven emotional evolution within a character remains underexplored. To address this limitation, we draw inspiration from the Component Process Model (CPM), a psychological theory that views emotion as a dynamic process shaped by the appraisal of external events. We propose CPM-MultiAgent, a CPM-grounded emotion evolution multi-agent framework for supporting emotional changes in persona-based dialogue. Instead of treating a character's emotion as a fixed attribute, CPM-MultiAgent represents it as a latent state that is continuously reshaped by dialogue triggers. Through affective trigger extraction, CPM-based collaborative appraisal, and emotion state updating, the framework enables more emotionally consistent role simulation in multi-turn interactions.Experiments with baseline comparisons, ablation studies, human evaluation, and case analyses demonstrate that CPM-MultiAgent effectively models dynamic emotional evolution in emotionally sensitive role-simulation settings.

View source

Similar papers

Jul 2026

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users'capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and organizes data, models, system design, evaluation, and governance around repeated use, non-use, transition, and termination. A targeted literature-and-corpus audit motivates this position. In a PRISMA-ScR-guided sample, 95% of 60 system-building papers pursue relief-oriented goals. None evaluates capability or longitudinal outcomes, and only 1 considers dependency, autonomy, or termination risk. In 300 ESConv supporter turns, capability-relevant functions appear in 43.0%, while generic suggestions account for 22.0%, compared with 4.0% reappraisal, 6.7% self-efficacy support, and 0.3% boundary behavior. We release a protocol for extending the audit to model behavior. An illustrative process model connects latent user capability to six design commitments, four evaluation timescales, and lifecycle constraints. The resulting agenda makes CSED testable across data, policy design, training, evaluation, and governance.

Ming Wang, Jiaqi Wu Young, Wenfang Wu et al. · 0 citations
Preprint Aug 2026

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their ability to capture embodied speaker--listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker's verbal and facial affective dynamics, estimates the robot listener's own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker--listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.

Z. Pang, C. Kennington, Tatsuya Kawahara · 0 citations
Open access 2026

Design and Evaluation of a Dual-Layer Emotion—Personality Framework for Adaptive Conversational Robots

—This paper presents a dual-layer emotional framework for human–robot conversational interaction that integrates internal emotion, representing the robot’s intrinsic affective state, and social emotion, representing outward emotional expression adapted for interpersonal alignment. Unlike conventional dialogue systems that rely primarily on semantic and contextual interpretation, the proposed framework processes user input through three complementary dimensions: content, context, and emotion, while supporting multimodal interaction through voice and physical actions. Emotion computation is regulated using two personality-modulated parameters. Sensitivity controls how strongly external stimuli influence internal emotional states, whereas consideration governs the degree to which the robot aligns its social expression with the user’s affect. A robotic prototype was developed to implement the framework, integrating multimodal sensors and expressive actuators for interactive operation. Preliminary experiments were conducted to evaluate the temporal evolution of emotional states, system response latency, and the influence of personality parameters on emotional behavior. The results illustrate that the system can update internal and social emotional states dynamically, adapt responses according to personality parameters, and generate emotionally coherent multimodal outputs. From a Human–Computer Interaction (HCI) perspective, the proposed framework provides a system-level approach for designing conversational interfaces capable of emotionally adaptive multimodal interaction in embodied robotic systems. The findings suggest that personality-modulated emotional regulation can support more flexible and context-dependent conversational behavior compared with conventional reactive emotional dialogue mechanisms.

Shitara Kaede, Kantawatchr Chaiprabha, Pimolkan Piankitrungreang et al. · 0 citations
Jul 2026

AtmosERC: Modeling Dialogue-Level Affective Atmosphere for Emotion Recognition in Conversation

Emotion Recognition in Conversation (ERC) aims to predict utterance-level emotions in dialogues and has largely advanced through context-centric modeling. However, global context is a heterogeneous signal, and not all contextual information is equally relevant to emotion prediction. This paper focuses on the affect-oriented component of this signal, termed dialogue-level affective atmosphere, which captures a latent tendency commonly reflected in conversational emotion patterns. To estimate and exploit this tendency, we propose AtmosERC, a graph-based ERC framework that models each dialogue as a conversational graph over utterances and speakers. A relation-aware graph extractor filters and fuses heterogeneous graph signals to produce dialogue-level and speaker-conditioned affective priors. The resulting compact prior guides lightweight sequential emotion prediction and can also be verbalized into prompt-level cues for LLM-based ERC without modifying backbone models. Experiments on four ERC benchmarks show that AtmosERC improves lightweight ERC, enhances LLM-based ERC as a plug-in cue, and yields more stable predictions under local emotional deviations.

Weijie Feng, Tong Zhang, Binbin Liu et al. · 0 citations
Book Open access Jul 2026

HCPRA: A Hierarchical Cognition–Perception–Reasoning Agent Framework for Emotion-Cause Pair Extraction in Conversations

Emotion-Cause Pair Extraction in Conversations (ECPEC) aims to identify speakers' emotions and the corresponding causes within conversation contexts. This task has gained increasing attention in recent years. Existing methods typically operate from the perspective of text object, inputting the entire conversation text and relying on neural networks to highlight utterance features as a means of emotion perception, while employing pair concatenation as a single path cause reasoning strategy. However, in conversation scenarios, humans are the true origin of emotions, while language text merely serves as the medium of expression. Consequently, these methods lack the modeling of the cognition and reasoning processes behind the human's behavior from the perspective of the speaker subject. To address this issue, we propose a novel Hierarchical Cognition–Perception–Reasoning Agent (HCPRA) framework. The framework is oriented around the speaker, simulating their personality cognition and conversation scenario cognition. Moreover, it utilizes hybrid memory and implicit emotion enhancement to simulate the human emotion perception. Additionally, it employs a hierarchical emotion cause reasoning mechanism to extract interpretable relationships between the speaker's emotions and their causes through emotion awareness and multi-path cause reasoning. Experimental results demonstrate that our approach achieves state-of-the-art performance on three datasets, with F1 improvement up to 15.61%.

Botao Wang, Lianwei Wu, Shuhan Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.