Aug 2026· 2 citations· ⚡ 1 influential· 31 references
Computer Science
TL;DR
CompanionBench, an interactive bilingual benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data, and is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data.
Abstract
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.
This paper advances a paradigm shift in AI alignment by treating personality not as an emergent side effect but as a primary, explicit training target. We present a rigorous data curation and supervised fine-tuning framework designed to produce an AI streamer that is both entertaining and agentic—capable of sustained i...
Jin-Sen Wu· AI and Data Science Journal· 0 citations
This Work-in-Progress examines whether personality-informed prompting changes perceived LLM emotional support in driving scenarios designed to elicit stress. A condition-order-balanced, within-subject CARLA simulator study (n = 14) compared a baseline with a Driver Personality Profile (DPP) condition; a supplementary o...
Max Mittelstädt, Ece Sutanrikulu, Lumbardh Ljatifi et al.· Adjunct Proceedings of the 1...· 0 citations
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified pro...
P. Fonseca, R. Rodríguez-Carvajal, Rafael A. Calvo· 0 citations
Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarch...
Rotem Dror, Zohar Elyoseph, Yuval Haber et al.· 0 citations
Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind...
BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.
Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.