Skip to content

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Jul 2026 · arXiv.org · Vol abs/2607.20536 · 1 citation · 31 references
Computer Science

TL;DR

AppWorld-UL is introduced, a ``user-in-the-loop''benchmark of 516 challenging tasks requiring diverse agent-user interactions that systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction.

Abstract

Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a ``user-in-the-loop''benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark's difficulty and its potential to advance research on user-in-the-loop tool-use agents.

View source

Similar papers

#natural language process... Preprint Aug 2026

PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems

This work introduces PersonaForge, a user simulation framework for synthesizing realistic multi-turn user--agent interactions that combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries.

Hanglong Lv, Dawei Zhu, Lei Li et al. · 0 citations
Jul 2026

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code extension, reliability remains below 70%.

Weihuang Zheng, Tianyuan Zou, Eileen Ye et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

GMA is presented, a benchmark for evaluating general mobile assistants in challenging real-world scenarios, and shows that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models.

Yi-Qi Zhu, Feiyu Gao, Jiakang Fan et al. · 0 citations
Jul 2026

ABISS: Evaluating Text-to-SQL Systems Through Agent Interaction

A unified taxonomy of 8 categories covering ambiguous and unanswerable questions is addressed, a multi-agent generation pipeline with a two-stage process (NLQ generation followed by SQL grounding) and an explicit Category Conformance validation stage are addressed.

Giovanni Sullutrone, Luca Sala, Sania Aftar et al. · 0 citations
Preprint Aug 2026

Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces

The findings suggest that the primary benefit of human--agent interfaces may be reducing interaction effort rather than improving speed, and that delegation reflects who the user is more than what the task demands.

Gavin Raine Dizon, Tyrone Justin Sta Maria, Jordan Aiko Deja et al. · 0 citations
Preprint Aug 2026

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

SWE-Touch is introduced, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to task-relevant code that conflict with task completion, and point to detecting workspace changes, reconciling conflicting edits with the task, and verifying the affected behavior as key capabilities for future optimization.

Yuqiao Tan, Jinxiang Meng, Fangyu Lei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.