A Multimodal Transformer-Based Digital Companion for Emotion-Aware Human–Computer Interaction Using Real-Time 3D Avatars
Abstract
Human-computer interaction has progressed significantly in recent years, yet existing digital assistants continue to suffer from limitations in emotional engagement, personalisation, and expressive communication. Most current AI systems operate primarily through text or voice, lacking a visual presence that allows users to form meaningful connections with technology. To address this gap, our paper introduces JARWIN (Just A Real-World Intelligent Network). This next-generation AI companion integrates a large language model, speech processing, and real-time 3D avatar animation to create a more natural, interactive, and intelligent user experience. The architecture features multiple operational tiers, including speech-to-text transcription, sentiment-aware conversational modelling, text-to-speech synthesis, and avatar-based emotional expression. Linear conversational modelling serves as the initial baseline for contextual understanding, while more advanced transformer-based models are employed to handle nuanced dialogue, sentiment variations, and continuous memory retention. Additionally, emotion recognition mechanisms dynamically map conversational tone to facial expressions and gestures, enabling human-like responsiveness. The result is a system capable not only of executing functional tasks, such as opening applications and retrieving information, but also of maintaining adaptive, emotionally aware dialogue. By bridging cognitive intelligence with visual embodiment, this paper demonstrates a scalable, innovative approach to immersive digital companionship, offering new opportunities for personalised assistance, interactive learning, mental wellness support, and human-AI relational computing.