Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· pp. 1693-1707· 0 citations· 25 references
Computer Science
TL;DR
This paper presents a service-aware network capacity planning suite that systematically improves network efficiency with a service collaboration approach, and proposes the "safe capacity" abstraction which enables services to incorporate current and projected network conditions into their compute and storage allocation decisions.
Abstract
Meta's rapid expansion in users, business operations, and AI workloads is straining our backbone network, while physical constraints—such as fiber, space, and power—limit the speed of capacity growth. To address these challenges, we present a service-aware network capacity planning suite that systematically improves network efficiency with a service collaboration approach. We propose the "safe capacity" abstraction which enables services to incorporate current and projected network conditions into their compute and storage allocation decisions. We introduce a hose-carving method that efficiently translates service-level traffic demands into detailed traffic matrices, allowing for more precise bandwidth allocation. To promote responsible network usage, we design a network rate card which attributes network consumption to individual services, incentivizing optimization and resource trade-offs. Additionally, new enforcement features at the end-host layer dynamically adjust resource allocations and traffic flows at runtime to maximize utilization. This paper is the first to detail a collaborative, service-aware approach to backbone network efficiency at Meta scale. Based on years of operational experience, we share practical insights and highlight new directions for research in network efficiency.
Compute and memory resources in cloud environments are strictly managed and isolated by the control plane; in contrast, network resources lack equivalent management and isolation mechanisms. This best-effort treatment of networking leads to significant challenges for modern AI workloads, which have diverse and bandwidth-intensive communication patterns. Without fine-grained network resource control, these workloads suffer from interference, unpredictable throughput, and suboptimal cluster utilization. To address these issues, this paper demonstrates how network bandwidth can be elevated to a first-class, schedulable, and enforceable resource within Kubernetes, the de facto standard for cloud-native orchestration. We introduce a new scheduling capability that models network interfaces as allocatable resources and regulates bandwidth sharing through the Dynamic Resource Allocation (DRA) framework, with enforcement implemented using the Hierarchical Token Bucket (HTB) mechanism. We evaluate the system using multitenant AI workloads derived from real-world communication characteristics with a simulation-based approach and validate the proposed enforcement strategy in a real cluster. Results show that the proposed two-level bandwidth allocation improves tenant performance predictability and satisfaction while maintaining packed cluster utilization.
This work proposes an enhanced Proximal Policy Optimization (PPO) framework for resource-aware and latency-sensitive SFC placement in edge-enabled networks, and demonstrates the applicability of the proposed framework in mission-critical and latency-sensitive service environments.
Nithin Melala Eshwarappa, Ching-Hsien Hsu, Hojjat Baghban et al.· ACM Transactions on Modeling...· 0 citations
This work introduces Online Pricing-based Slice Admission Control and Resource Allocation (OPA) framework, which dynamically assigns pseudo-prices to resources that capture long-term scarcity and anticipated inter-temporal opportunity costs and designs an exponential pricing strategy that guarantees bounded worst-case performance.
Muhammad Sulaiman, Bo Sun, M. A. Salahuddin et al.· 0 citations
A constrained optimization model that supports different management goals through alternative objective functions (latency-aware or power-aware) while enforcing operational constraints, including node capacities, slice-specific latency bounds, and explicit limits on VNF migrations/relocations between scheduling periods is proposed.
R. Moreno-Vozmediano, E. Huedo, R. Montero et al.· Journal of Network and Syste...· 0 citations
This work proposes CAPSUM, a capacity-aware admission policy with an elastic specialization, CAPSUM-E, and implements an exact local offline dynamic program and compares against direct common-model baselines and documented source-derived adapters for EDP-A, OREO, and uEDC-L.
Hailiang Zhao, Ziqi Wang, Yi-Fei Zhang et al.· 0 citations
The results show that soft SLO limits reduce corrective rescheduling actions by 49% compared to hard-limit approaches while maintaining acceptable performance guarantees, and resource-aware scheduling decreases node-level congestion and further mitigates SLO violations, demonstrating the effectiveness of incorporating application-level flexibility and hardware-level insights into scheduling and rescheduling decisions.
Oliver Larsson, Thijs Metsch, Cristian Klein et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.