A reference architecture is developed that separates proposal generation from governed execution, identifies recurring integration patterns for LLM-enabled operations, and derives a research agenda for higher network autonomy under explicit assurance, safety, and governance constraints.
Abstract
Since the inception of modern communication networks, the quest for operations automation has never ceased. Yet the evolution of network automation is difficult to characterize with a single maturity ladder. Throughout this history, network control systems have expanded their capabilities for observation, decision support, routine execution, and operator interaction, but these capabilities have not advanced uniformly. Such uneven progress makes the degree of automation an unreliable proxy for trustworthy network-side actuation. The unresolved question is not simply how much automation a system provides, but under what conditions it can be entrusted to change the network state. This paper examines that question through Network Control Intelligence (NCI), a five-axis framework spanning Decision Logic, Adaptability, Knowledge, Control Delegation, and Interface. We use NCI to organize the evolution of network-control systems into three eras: rule-based and scripted automation, programmable and data-driven control, and Large Language Model (LLM)-enabled network operations. Viewed through this framework, the three eras reveal a persistent asymmetry. None of these gains, however, automatically determines when network control should be trusted to change the network state. We frame trustworthy autonomy as a governed alignment between what a system can infer, what it can verify, and what it is authorized to execute. On that basis, the paper develops a reference architecture that separates proposal generation from governed execution, identifies recurring integration patterns for LLM-enabled operations, and derives a research agenda for higher network autonomy under explicit assurance, safety, and governance constraints.
The study provides initial evidence of feasibility while identifying the challenges that must be addressed before production deployment and formalize the ADN agent model and workflow and define an operational framework covering communication, lifecycle management, governance, and security.
F. Rossi, Paulo Silas S. De Souza, Diogo M. Monteiro et al.· IEEE Access· 0 citations
Agentic artificial intelligence is increasingly integrated into communication networks, enabling systems that perceive network state, reason over conditions, and execute actions on live infrastructure. As these systems act directly on operational environments, their behavior after failure remains insufficiently characterized: existing evaluation emphasizes throughput, latency, and accuracy, without capturing how systems recover after actions that affect network state. This survey examines agentic artificial intelligence in communication networks through the lens of recoverability, based on a systematically coded corpus of the recent literature. The paper introduces the MATR-R framework, which organizes recoverability around memory awareness, action governance, trust regulation, and recovery capability. At its core lies the R0–R5 graded recovery scale, which classifies post-failure capability on an ordinal scale ranging from detection to learning. The survey further presents a taxonomy of communication-network failure modes, a recovery-aware evaluation framework, and five tutorial scenarios across slicing, cybersecurity, and digital-twin-assisted recovery. The analysis reveals three consistent patterns: containment is the most common recovery level, self-healing is more often claimed than demonstrated, and full autonomy does not coincide with recovery beyond containment. These findings highlight a gap between perceived and verified recovery capability. Recoverability emerges as a distinct design dimension for agentic communication networks, with open challenges in recovery benchmarking, post-recovery verification, and standardization.
Omar K. Dawoud, Osama A. Ghoneim· Sustainable Machine Intellig...· 0 citations
Operating complex communication networks (NWs) requires rapid, reliable failure recovery, yet fully autonomous recovery in realistic settings remains elusive. We present a failure-recovery-specialized LLM agent that encodes the operator’s workflow into a structured reasoning-and-acting process. A core design is stage-constrained operation, which regulates tool admissibility by operational stage to prevent unsafe actions and maintain closed-loop stability. The agent interacts with a device-accurate NW digital twin built on Cisco Modeling Labs via diagnostic and control tool calls under these constraints. In evaluations on diverse interface (IF) and link failure patterns, the framework enabled the tested commercial LLMs to autonomously recover most failures, with frontier-class models achieving full recovery in every setting. Augmenting diagnostic outputs with baseline diffs improved accuracy across all models, and command-aware retrieval-augmented generation (RAG) further raised success for cost-effective models by supplying proven precedents. Case studies further show fully autonomous handling of complex failures such as IF flapping and a configuration-induced routing loop, demonstrating that a structured agent framework can bridge LLM reasoning and the demands of practical NW failure recovery.
Hiroki Ikeuchi, Yousuke Takahashi, Hitoshi Shimizu et al.· International Conference on...· 0 citations
AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state. Nonetheless, operational networks usually span many devices and administrative domains. Realizing an operator's intent requires coordinating agents with distinct authority scopes that define the resources they can access, the operations they can invoke, and the network state they can observe. This division limits the blast radius of an erroneous action but fragments the evidence needed to assess the network-wide outcome. Successful execution of a configuration action proposed by one agent does not establish that remote devices responded as intended or that routing changes reached the required devices. A valid observation may also become stale after a subsequent change. Before the coordinated operation can be declared complete, a trusted assurance layer must collect current observations from the required scopes and determine whether they collectively support the operator's intended network-wide outcome. To address the completion admission problem, we present EvidenceNet, a runtime assurance layer for deciding whether coordinated agent operations have achieved an operator's network intent. Its broker collects the post-change observations required by a completion contract, and its admission gate checks that the evidence comes from the required scopes, remains current, and satisfies the task rules. A verifier agent provides an additional assessment of the observation content. Experiments on live routing networks show that post-change state checks recognize successful outcomes that configuration-action records alone cannot establish. Controlled interventions further show that EvidenceNet rejects completion when otherwise satisfactory observations have the wrong source, have been substituted, or are stale.
This work proposes a new logically centralized controller powered entirely by an LLM, called Agentic-Defined Networking (ADN), a novel architecture that integrates LLMs as the reasoning core of an SDN control plane implemented in a real network controller.
Shanaya Varkey, Sean Choi· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.