Do Prompt-Level Empathy Instructions Influence User Experience? Evidence From A Controlled Chatbot Study
Empathic behavior is widely considered a key design goal for conversational agents, and prompt engineering is often used to shape empathic interaction styles in large language model (LLM) systems. However, it remains unclear whether prompt-level empathic framing produces measurable differences in user experience. We conducted a controlled between-subjects study comparing two chatbot configurations differing only in system-level prompts: a neutral assistant and an empathically framed variant inspired by Affect Control Theory. Participants engaged in conversations about exam anxiety and completed standardized measures of affective state (PANAS), perceived empathy (PETS), and usability (CUQ); an automated manipulation check was run in parallel. Across perceived empathy, usability, and response-level empathy ratings, observed differences between conditions were small and statistically inconclusive, while affective change estimates pointed in the direction predicted by empathic framing without reaching significance at this sample size. The paper’s primary contributions are methodological: a controlled isolation of prompt-only behavioral control in aligned LLMs, evidence that automated response-level and user-perceived empathy can dissociate, and a case for multi-level evaluation strategies—response-level, perceptual, and affective—when assessing empathic conversational systems.