EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents
Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchma...