Preprint
Aug 2026
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
AGENTCHAOSBENCH is presented, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry, and its held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.
Chenkai Zhang, Yiran Li, Yifang Tian et al.
· 0 citations