Skip to content
Conference

A Multimodal Emotional Interaction Framework Driven by Large Language Models

Jul 2026 · 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM) · pp. 1-7 · 0 citations · 18 references

Abstract

Natural and empathetic human-robot interaction is essential for social robots and other AI applications, while emotional feedback is less discussed. Thus, this paper proposes a multimodal interaction system driven by large language model (LLM). The system constructs a unified emotional state vector by integrating visual (facial expressions) and auditory (speech emotion) cues. It employs DeepSeek LLM for context-aware chain-of-thought reasoning to generate contextually appropriate verbal responses, facial expressions, and head movement commands. To achieve optimized latency and fluid embodied interaction, the system adopts a layered architecture: the upper-level LLM handles semantic understanding and behavior planning, outputting structured JSON commands; the lower-level controller generates smooth motion trajectories and manages multimodal interaction flows using a finite state machine (FSM). Experimental results demonstrate the system’s ability to effectively resolve emotional ambiguities, track emotional evolution during continuous dialogue, and achieve an optimized end-to-end response latency. This validates its feasibility and engineering value in practical human-robot interaction scenarios.

View source

Similar papers

Open access 2026

Design and Evaluation of a Dual-Layer Emotion—Personality Framework for Adaptive Conversational Robots

—This paper presents a dual-layer emotional framework for human–robot conversational interaction that integrates internal emotion, representing the robot’s intrinsic affective state, and social emotion, representing outward emotional expression adapted for interpersonal alignment. Unlike conventional dialogue systems t...

Shitara Kaede, Kantawatchr Chaiprabha, Pimolkan Piankitrungreang et al. · 0 citations
#human-computer interacti... Preprint Aug 2026

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

This paper introduces OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction, and proposes Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation.

Jiahao Huang, Zheng Lian, Jingyi Zhang et al. · 0 citations
Open access Aug 2026

A Multimodal Transformer-Based Digital Companion for Emotion-Aware Human–Computer Interaction Using Real-Time 3D Avatars

Human-computer interaction has progressed significantly in recent years, yet existing digital assistants continue to suffer from limitations in emotional engagement, personalisation, and expressive communication. Most current AI systems operate primarily through text or voice, lacking a visual presence that allows user...

A. S. Sundhar, J. A. Jeba, V. Ramkumar et al. · 0 citations
Open access Aug 2026

Large Language Model-Assisted Distillation–Fusion Framework for Visual Emotion Recognition

A large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability and adaptively fuses predictions from the instruction-tuned VLM and the distillation modul...

Yujun Ma, Yun-Jie Zeng, Zhi-Yuan Chen et al. · 0 citations
Preprint Aug 2026

Aura: Dynamic Intra-Turn Emotion-Aware Adaptation of Large Language Model Responses

Effective human-AI interaction requires systems that dynamically adapt to a user's behavior and evolving understanding. When users interact with Large Language Models (LLMs), these models typically respond to prompts without sensing the user's immediate reactions. This lack of communicative synchrony can lead to inform...

Rachel Schuchert, Christian Holz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.