Skip to content
Open access

Large Language Model-Assisted Distillation–Fusion Framework for Visual Emotion Recognition

Aug 2026 · Algorithms · 0 citations · 23 references

Abstract

Visual emotion recognition plays a critical role in human–computer interaction and mental health applications. Although existing Vision–Language Models (VLMs) alleviate the limitations of conventional vision models in high-level semantic understanding, they still face three main challenges: limited emotional semantic understanding, insufficient visual emotional perception capability, and high computational costs when deploying both models simultaneously. To address these issues, a large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability. Furthermore, we transfer the visual emotion discrimination knowledge of a conventional vision model into the VLM using a distillation module while keeping the VLM frozen during training, which reduces the computational costs. Following that, we design a fusion and prediction module that adaptively fuses predictions from the instruction-tuned VLM and the distillation module for final emotion recognition. The experimental results on the Abstract, ArtPhoto, Emotion6, and FI datasets demonstrate that VERLADF achieves recognition accuracies of 36.71%, 52.38%, 74.73%, and 79.69%, respectively, significantly outperforming many methods in the literature and demonstrating the effectiveness of the proposed framework.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.