Towards Emotional Intelligence in Conversational AI: How Well Can LLMs Recognise Emotion in Conversations?
The ability to express contextually appropriate emotions remains a defining challenge in conversational AI. Meeting this challenge requires accurately annotated corpora, a need that inevitably collides with practical constraints. Manual annotation, while reliable, is often prohibitively expensive and time-consuming. Consequently, automatic emotion recognition has become a critical component in the fine-tuning pipeline, with Large Language Models (LLMs) emerging as a promising alternative to human annotators. Despite their considerable potential for generating training and evaluation data, the use of LLMs for this purpose introduces a subtle but significant risk: the creation of circular, potentially biased evaluation protocols that frequently go unacknowledged in empirical reporting. In this paper, we present a systematic evaluation of LLMs for conversational emotion recognition, using human annotation as a gold standard to quantify the degree to which reliance on LLMs may introduce error into both fine-tuning and evaluation workflows. We conduct a comprehensive assessment of multiple LLMs for emotion labeling across conversational contexts, and additionally examine a second critical factor: the impact of conversational context on LLM performance. Our results indicate that although LLMs benefit from access to prior conversational context, their utilization of such context differs substantially from human conversational understanding, suggesting that LLMs may not rely on the temporal ordering of conversational context in the same way humans do.