Skip to content
Preprint

ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

ADeptS-Bench is introduced, a dual-stream trustworthiness benchmark, grounded in the ADEPTS capability framework and general population user studies, and reveals that all models overestimate consequence severity, mirroring the over-refusal bias observed in safety.

Abstract

Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can safely interact with visual interfaces while handling ambiguous instructions. We introduce ADeptS-Bench, a dual-stream trustworthiness benchmark, grounded in the ADEPTS capability framework and general population user studies. The Safety stream provides paired benign/malicious tasks with threats embedded in the visual interface. The Disambiguation stream evaluates whether agents seek clarification when intent is ambiguous. Evaluating seven models reveals that no model consistently exceeds 80% task success while staying below 30% attack success; every model clicks"Checkout"on a $25K order without hesitation, and none detects that a"factory reset"button is mislabeled as"Optimize."An ablation reveals three distinct safety architectures: tool-dependent (ASR +21-23pp without refusal tool), partially tool-dependent (+10-11pp), and no mechanism (unchanged). In disambiguation, all models overestimate consequence severity, mirroring the over-refusal bias observed in safety. We release all data, evaluation code, and analysis tools upon publication.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

SIR: Self-improving Red-teaming for Compute Use Agents

SIR is presented, a black box IPI attack that composes stealthy injections from a small library of reusable principles stated in plain language and wraps composition in an iterative feedback loop that diagnoses the victim's failed trajectories and distills the bypasses into new, named strategies that are reapplied acro...

Chen Xiong, Zhi-Yuan He, Pin-Yu Chen et al. · 0 citations
Jul 2026

How Benchmarks Mis-Score Computer-Use Agents

A three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain, and connects these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.

Zi-Han Dong, Zhiyuan Ma, Zekun Wang et al. · 2 citations
Jul 2026

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion

It is shown that a one-line guardrail achieves large single-shot ASR reductions, up to roughly 40 points, at near-zero over-refusal cost, which overstates deployed robustness by a systematic and predictable margin.

Haoxin An, Yunpeng Song, Zihao Bai et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CURA: Certified Runtime Alarms for Computer-Use Agents

This work introduces CURA (Certified Runtime Alarms for Computer-Use Agents), an external monitor that reads only harness-visible telemetry, with no model internals, extra LLM calls, or prompt changes, and turns the running trajectory into a sequential test with certified false-alarm control.

Divake Kumar, Sina Tayebati, Devashri Naik et al. · 0 citations
Jul 2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.

Qiushi Sun, Kanzhi Cheng, Yian Wang et al. · 1 citation · ⚡1
Conference 2026

When Verification Hurts: The Cost of Overriding Abstention in Two-Stage Web Agents

This study cautions against transplanting verification into grounding pipelines and identifies calibrated abstention as a property worth preserving and proposes an abstention-aware verifier that intervenes only under sufficient candidate coverage and confidence.

Duchen Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.