Skip to content
Preprint

Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

The findings suggest that the primary benefit of human--agent interfaces may be reducing interaction effort rather than improving speed, and that delegation reflects who the user is more than what the task demands.

Abstract

Large Language Models (LLMs) are increasingly embedded into applications, allowing users to complete tasks either through direct manipulation or by delegating actions to conversational agents. However, little is known about how users balance these modalities when both are available. We present a web-based content management system augmented with an LLM agent through the Model Context Protocol (MCP), enabling users to perform CRUD tasks through a graphical interface, a conversational agent, or both. We conducted a between-subjects study (N=73) comparing three interaction modes: Traditional-Only, AI-First, and Hybrid. Across sixteen scenarios, we analyzed task completion time, interaction logs, and delegation behavior. AI-assisted interaction significantly reduced clicks, page navigations, and scrolling indicating lower interaction effort. Surprisingly, these reductions did not translate into faster task completion, as task duration did not differ significantly across conditions. We also found no significant relationship between CRUD operation type and delegation, suggesting that users did not systematically avoid delegating higher-risk actions. Instead, delegation varied far more between participants than between tasks, with individual differences accounting for roughly half the variance in assistant use (ICC = .50). Our findings suggest that the primary benefit of human--agent interfaces may be reducing interaction effort rather than improving speed, and that delegation reflects who the user is more than what the task demands.

View source

Similar papers

Jul 2026

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

AppWorld-UL is introduced, a ``user-in-the-loop''benchmark of 516 challenging tasks requiring diverse agent-user interactions that systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction.

Junzhi Chen, H. Trivedi, Jane Pan et al. · 1 citation
Preprint Jul 2026

Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents

Computer Use Agents (CUAs) can autonomously execute complex, multi-step tasks within GUIs, enhancing efficiency through parallel multitasking. However, our formative studies with CUA experts and GenAI users indicated that current feedback is primarily text-based, requiring sustained attention to monitor progress and offering limited visibility to trace past GUI interactions. Based on the findings, we developed a prototype system, Sidekick, for communicating CUAs'status with multimodal feedback across different stages of interaction: (i) When CUAs run in the background, Sidekick signals its execution state through ambient cues. (ii) Upon resuming interaction with CUAs, Sidekick provides multimodal summaries of completed actions to support rapid context resumption. (iii) When CUAs operate in the foreground, Sidekick enhances transparency by verbalizing and visualizing the agent's reasoning. A study with 30 participants demonstrated that Sidekick significantly improved multitasking performance with CUAs compared to baseline systems that presented textual feedback either in a typical chat or in an ambient display. Sidekick supported progress awareness, and error and action traceability more effectively. Finally, we demonstrate the promise of Sidekick through several example applications, and discuss implications for long-horizon human-agent collaboration.

Ruei-Che Chang, Wenqian Xu, Dingzeyu Li et al. · 0 citations
Jul 2026

Just A Rather Very Intelligent Spoken Agent

JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows, is introduced and preliminary results suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments.

Chen Chen, Zhehuai Chen · 0 citations
Book Open access Jul 2026

Who is Behind the Voice? User Reasoning About Agency and Capability in a Multi-Party Conversational Interface

The potential for conversational user interfaces (CUIs) to collaborate and even lead human teams engaged in a collaborative activity remains an intriguing yet largely underexplored application of CUIs. Specifically, little seems to be understood about how people perceive such agents and their reasoning behind it. To begin to examine this intriguing topic, we performed a thematic analysis on open ended questions posed to participants after being led by either a human or CUI based leader on a collaborative task. Analysis revealed five key themes around how they perceived the leader they were interacting with: guidance, voice, understanding, timing and behaviour. These themes ultimately shaped how participants reasoned about whether the leader guiding them was a human, or autonomously controlled, highlighting key considerations designers should consider when creating such collaborative conversational agents.

James Simpson, Hamish Stening, Gaurav Patil et al. · 0 citations

A Goal-Oriented Agentic Framework For Collaborative Branching Human-Robot Interactions

It is suggested that the benefit of the agentic framework lies primarily in interaction quality rather than conversational efficiency, and the agentic architecture displayed robustness by recovering from non-normative inputs while maintaining strict goal alignment.

Morten Roed Frederiksen · 1 citation
#artificial intelligence Preprint Sep 2026

Are We There Yet? Assessing Computer-Use Agents for Blind Users'Accessible Interaction with Desktop Applications

Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows remains unclear. We present a three-week diary study with 8 blind users using OLLA, a screen-reader-accessible CUA prototype, collecting 1,258 commands across 12 applications with screenshots, UI trees, model responses, and action traces. We evaluate GPT-5 during deployment and re-execute the same commands with four additional models. GPT-5 achieved the highest success rate at 52.5%. Trace analysis reveals grounding, planning, constraint-tracking, and termination failures, while interviews reveal beyond-automation needs.

Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.