Emotion Recognition in Conversation (ERC) requires models to identify subtle emotional cues that are often distributed across distant dialogue turns. Existing methods typically incorporate dialogue history through a fixed context window. However, short windows discard potentially useful long-range evidence, while enlarging the window repeatedly re-encodes overlapping utterances, increases computational and memory costs, and may introduce irrelevant context. Moreover, commonly used parameter-efficient adaptation methods, such as LoRA, mainly introduce fixed low-rank transformations in the feature space and do not explicitly maintain a dialogue-level state or condition their transformations on the evolving conversational context. To address these limitations, we propose a lightweight adapter, DiaRelay, to enable LLMs to explicitly maintain a dialogue-level memory for accurate ERC. Based on LoRA, DiaRelay introduces two extra tightly collaborative components, Selective Relay Memory Transition and Dual-axis Relay Memory Read. Selective Relay Memory Transition progressively aggregates useful historical evidence into a bounded relay memory and propagates it across successive utterance predictions. This allows earlier emotional cues to influence later predictions after they leave the local context window, without re-encoding the complete dialogue history or expanding the backbone context length. Dual-axis Relay Memory Read uses the propagated memory to dynamically modulate low-rank feature transformations, enabling context-dependent representation adaptation without test-time gradient updates. Extensive experiments show that DiaRelay can achieve SOTA weighted F1 and accuracy on MELD while obtaining competitive results on IEMOCAP with only an extra 7.1M trainable parameters, indicating the effectiveness and generalizability of our DiaRelay in enhancing LLM-based emotional understanding.
Zi-Hao Zhou, Bin Yang, Jinghui Qin et al.· 0 citations
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.
Lihuang Fang, Yuchen Zou, Ke-Bing Jin et al.· arXiv.org· 1 citation
Low-altitude embodied intelligence (LAEI) has emerged as a promising solution for operational efficiency and sustainability of the emerging low-altitude economy via perception–reasoning–action loops. The ground base stations with limited service coverages fail to achieve ubiquitous connectivity in widespread environments. The aerial agents embedded in flying bodies ensure pervasive intelligence across dynamic three-dimensional spaces. However, the joint optimization of flight trajectories and resource allocation for hierarchical UAV networks introduces large state and action spaces, posing significant challenges for real-time mission execution. In this paper, we propose an agentic Generative Artificial Intelligence (GenAI)-based LAEI framework. In the framework, a joint optimization problem is formulated to minimize long-term average energy consumption while ensuring task queue stability and satisfying spatial kinematic constraints. The Lyapunov optimization technique decomposes the long-term energy minimization problem into deterministic per-slot sub-problems with low computational complexity. A diffusion-based GenAI algorithm synthesizes optimal trajectories through an iterative denoising process, where the model-based resource allocation problem serves as guidance to accelerate convergence. Finally, extensive simulation experiments indicate that the proposed GenAI-enabled algorithm outperforms other baseline schemes, delivering minimized energy consumption and enhanced resource utilization in dynamic low-altitude embodied intelligence environments.
Dong-Hai Wu, Jiangtian Nie, Yang Zhang et al.· IEEE Transactions on Cogniti...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.