CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
CAP is introduced, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding and a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, complex execution operations, and perceptual requirements, and then recomposes these components into realistic cross-site workflows.