This work presents SRE-Marathon, a benchmark for long-horizon, continuous SRE operation, where an agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, p...
Yi-Fang Tian, Ying-Jian Bai, Yi-Feng He et al.· 0 citations
Results show that operational telemetry can be transformed from diagnostic evidence into actionable repair context: paired telemetry supports repair-oriented localization, while repair graph agents convert localized code and configuration evidence into constrained patch-generation context for the LLM.
Yuan-Chen Gao, Yi-Fang Tian, Yi-Ran Li et al.· 1 citation
AGENTCHAOSBENCH is presented, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry, and its held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.
Chenkai Zhang, Yiran Li, Yifang Tian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.