Skip to content

UXBench: Benchmarking User Experience in AI Assistants

Jun 2026 · arXiv.org · Vol abs/2606.09570 · 1 citation
Computer Science

TL;DR

The first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation is presented, demonstrating that user feedback prediction is a learnable capability and revealing different aspects that influence user experience.

Abstract

As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present \textbf{UXBench}, the first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation. The benchmark consists of three interconnected tasks, UX Judge, UX Eval, and UX Recovery, with 7,400 test instances extracted from over 70K interaction logs of a mainstream Chinese AI assistant. The dataset closely reflects real user distributions, covering 8 scenarios, 83 domains, and diverse failure patterns that pose severe challenges. Extensive experiments on 26 frontier language models provide novel insights into how well models perceive user experience and how improvements in model capability contribute to better dialogue engagement. Through analyses of model behavior and performance gaps, we document six important findings, demonstrating that user feedback prediction is a learnable capability and revealing different aspects that influence user experience. UXBench establishes a new evaluation landscape and calls for greater attention to tailored UX optimization, contributing toward a user-centric scaling law for the development of successful AI assistants. The full project is released at https://github.com/mengze-hong/UXBench.

View source

Similar papers

#artificial intelligence Review Sep 2026

Automatic multimodal UX improvement recommendations from LLM agent user simulations

Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manu...

Anurupa Chowdhury, Bin Wu, Hossein A. Rahmani et al. · 0 citations
Preprint Aug 2026

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...

Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al. · 0 citations
Preprint Oct 2026

Beyond Task Completion: Measuring Interaction Cost in Terminal User Interfaces

Large language models (LLMs) are increasingly used through terminal user interfaces (TUIs), yet task completion alone does not capture how difficult an interface is to understand and operate. Existing human assessments and LLM-generated ratings or reports do not provide repeatable measurements of interaction effort gro...

Rui-Da Hu, Yuan-Hao Wang, Chao Peng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Large-Scale User Behavior Analysis in Multimodal AI-Assisted Manual Task Execution

Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world tasks such as cooking and DIY through voice, text, image, and video interactions. Prior user studies have focused on controlled settings, leaving limited understanding of real-world CTA usage at scale. In this...

Rafael Ferreira, Diogo Tavares, Diogo Glória-Silva et al. · 0 citations
#artificial intelligence Preprint Sep 2026

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

This work introduces AgenticInterleave, a single-agent ReAct framework for retrieval-supported answer generation, together with IVR-12, a 12-dimensional rubric for assessing the content, presentation, and image quality of interleaved references and model outputs, and evaluates both short-answer correctness and interlea...

Hao-Nan Jiang, Guo-Jian Zhan, Jian-Cong Xie et al. · 0 citations
Preprint Aug 2026

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

This work introduces UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive, and establishes per-turn intent control as a complementary dimension to response fidelity in user simulation.

Bo Wang, Ruixing Zhang, Yunqi Liu et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 8, 2026

Flint: A visualization language for the AI era

Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.