DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedback, is introduced, an executable benchmark for agentic lifecycle control under delayed feedback that measures detection utility with average precision and capture at fixed review depth, and characterize adaptation through model switches and compute.
Abstract
Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedback. It comprises seven temporal graph datasets with node- and edge-level anomaly tasks, eleven selectable detectors, and eight chronological deployment windows per dataset. In each window, a controller sees only time-causal aggregate context, registered model cards, and its own matured history. It must choose a detector before current-window training or candidate scores exist. A sandboxed executor trains the chosen architecture on mature data, scores a hidden deployment window, and releases the outcome after a one-window delay. A deterministic verifier checks decision timing, leakage guards, legal actions, training scope, and persisted artifacts. We measure detection utility with average precision and capture at fixed review depth, and characterize adaptation through model switches and compute. Complete eight-window trajectories from two primary controllers and a no-memory reference on four datasets, together with three additional controllers on three datasets, expose useful, costly, and ineffective reactions to delayed evidence without granting an exhaustive current-window oracle.
Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
Seonglae Cho, F. Fernandez, Umar Mohammed et al.· 1 citation
The result is a bounded rule for interpreting aggregate agent behavior: first establish exposure, then score change, and abstain when the trace cannot support the claim.
These findings show why successful outcome recovery is insufficient evidence of exactly-once execution, and establish conformance within one calibrated deterministic testbed, rather than general safety or real-world self-improvement.
An extensive empirical study of state-of-the-art DGAD models is conducted, revealing that detection difficulty increases consistently with anomaly complexity, from simple localized irregularities to coordinated and temporally persistent structures.
Mohamed Nazim Mezhoudi, Guillaume Lachaud, Yan-Lei Diao et al.· Proceedings of the 32nd ACM...· 0 citations
SimpleCount is proposed, a reference with no parameter fitting that selects one scalar feature per dataset from a fixed pool of counts, recencies, first-occurrence indicators, and count-derived transforms that matches or exceeds SLADE on three of six datasets and exceeds IsoForest on all six.
It is argued that anomaly detection for agentic AI must reason at the workflow level, where global execution structure exposes signals that local checks cannot see, and presents Skynet, a principled workflow-level anomaly detection framework that turns observed multi-agent execution into directed workflow graphs and sc...
Chao-Yu Zhang, He-Xuan Yu, Heng Jin et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.