The results support conformal shielding as a practical mechanism for exposing and controlling the safety-speed trade-off in learning-based routing, while also identifying the limits of guarantees under topology and load shift.
Abstract
Adaptive routing controllers based on deep reinforcement learning can shorten recovery after failures, but a policy optimized only for expected reward may select timer and damping configurations whose control overhead or route oscillation risk is poorly estimated under rare or shifted conditions. This paper presents ConformalSafe-Routing, a runtime safety layer for adaptive routing convergence. A dueling double deep Q-network ranks bounded routing profiles, while an independently trained risk model estimates one-step operational cost from topology, failure, load, and convergence-state features. Split conformal calibration converts point predictions into finite-sample upper bounds. A horizon-aware allocation uses a per-decision miscoverage budget of 0.0125 for an eight-step episode, and the shield selects the highest-value action whose upper bound satisfies the safety envelope; a balanced static profile is used when the certified set is empty. Experiments use four real telecom topologies from the Internet Topology Zoo and a fully released topology-driven event simulator. Across 300 paired episodes per condition, the proposed method reduces episode-level safety violations from 56.7% to 4.3% in-domain and from 64.0% to 2.0% under high-load distribution shift relative to unshielded DQN. It also reduces route flaps by 73.0% and 76.6%, respectively. Compared with a balanced static profile, it shortens mean convergence by 22.5% in-domain and 9.9% under shift while maintaining low violation rates. The results support conformal shielding as a practical mechanism for exposing and controlling the safety-speed trade-off in learning-based routing, while also identifying the limits of guarantees under topology and load shift.
A reproducible discrete-time link-state convergence proxy with independent topology views, stochastic failure detection and link-state advertisement diffusion, shortest-path recomputation, and forwarding-loop/drop checks is implemented.
David Clarke, Mei Huang, Jonas Eriksen· International journal of inf...· 0 citations
Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remains an open challenge. Two obstacles limit the adoption of learning-based network control in this setting: traditional volume-based routing metrics, while highly effective for general traffic management, are not designed to capture traffic urgency; and Deep Reinforcement Learning (DRL) controllers trained from scratch suffer from sample inefficiency, long training times, and early-stage exploration volatility. This paper introduces a deployment-focused network control framework that addresses both obstacles. First, we present Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into Multi-Agent Deep Reinforcement Learning Effective Congestion ($p^*$) (MADRL EC ($p^*$)), a hybrid architecture combining a distributed scheduler with a centralized RL-based router. Second, we introduce a unified training objective that generalizes existing policy-learning paradigms---behavioral cloning, offline Reinforcement Learning (RL), online RL, and offline-to-online schemes---as special cases, combining a live-reward term, a pre-collected-reward term, and a policy-imitation term. From this objective, we derive the Model-Guided Annealed Reinforcement Learning (MGA-RL) protocol, instantiated on a Deep Deterministic Policy Gradient (DDPG) backbone: a deployment-oriented, demonstration-driven training approach that generalizes conventional Offline-to-Online (O2O) schemes, in which trajectories from a lightweight [...]
Vincenzo Norman Vitale, Mohammad Solki, A. Tulino et al.· 0 citations
GraphRoute-Transfer, a graph-neural-network policy that assigns per-node timers from local structural features and is by construction permutation- and size-invariant, is proposed, a graph-neural-network policy that assigns per-node timers from local structural features and is by construction permutation- and size-invariant.
Yuto Nakamura· Journal of Computing and Ele...· 0 citations
GraphRoute-Transfer, a graph-neural-network policy that assigns per-node timers from local structural features and is by construction permutation-and size-invariant, is proposed, a graph-neural-network policy that assigns per-node timers from local structural features and is by construction permutation-and size-invariant.
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al.· Technologies· 0 citations
This work proposes Double-Channel Graph Attention (DCGA), an end-to-end reinforcement learning framework that isolates network reachability and demand-service logic into separate graph channels and constructs valid routes using a simulator-coupled, constraint-informed decoder.
Hao Sun, Fang He, Congyuan Ji et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.