This work presents SRE-Marathon, a benchmark for long-horizon, continuous SRE operation, where an agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, p...
Yi-Fang Tian, Ying-Jian Bai, Yi-Feng He et al.· 0 citations
Lara, a machine-checkable language and protocol for checking and revising support for research claims, is introduced, and the metatheory of claim checking and cross-context argument transport is established, and semantic guarantees in Lean 4 are mechanized.
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Sheng-Han Zheng, Zong-Lin Di, Yimin Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.