We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...
Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al.· 0 citations
Recommendation systems thrive on personalization, where “correctness” is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality guidelines, they face a distinctive challenge: context-aware preference alignment. Recent...
Jun-Cheng Dong, Ding Tong, Ishan Gupta et al.· Proceedings of the 20th ACM...· 0 citations
This work argues that an LLM judge running in a production system is better understood as having a lifecycle: it must be built, trained, deployed, and continuously maintained as the surrounding data evolves, and each phase poses distinct technical and operational challenges.
Emma Kong, J. Tan, Ishan Gupta et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.