2026· IEEE Transactions on Network Science and Engineering· Vol 13, pp. 11513-11531· 0 citations· 78 references
Abstract
Deploying trillion-parameter large language models across metropolitan environments is required to sustain real-time inference. Urban power constraints, however, prohibit monolithic GPU clusters, forcing the integration of distributed supernodes into a citywide compute fabric. Over 100-km distances, optical propagation delays invalidate reactive congestion control for 12.8 Tbps cross-domain pipeline parallelism (PP) traffic. The resulting high bandwidth-delay product causes in-flight data to saturate edge switch buffers before the feedback loop triggers source throttling. We address this limitation with M-CSN, an AI compute architecture featuring LOTUS, a model-driven scheduling engine for metropolitan computing power networks (CPNs). LOTUS replaces delayed reactive signaling with a proactive bi-level scheduling strategy. At the macro-layer, an inverse-SLA mechanism partitions capacity among concurrent flows to guarantee performance isolation. At the micro-layer, the scheduler decouples transmission rates from dynamic window probing: it injects calculated pacing rates and static hyper-BDP windows into source RDMA queue pairs (QPs). This source-side enforcement eliminates feedback lag. Packet-level simulations demonstrate that LOTUS sustains a 92% operational load with PFC-free execution. Eliminating pause-induced jitter and buffer saturation reduces critical-flow latency by 2.6× relative to reactive baselines. This enables fragmented urban resources to operate as a unified, deterministic, supernode-based computing power network.
This work introduces CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header, and proposes Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production.
Abhiram Ravi, Nandita Dukkipati, Weiwu Pang et al.· Conference on Applications,...· 0 citations
The nascTime framework on OMNeT++/Simu5G is used to evaluate how many TSN endpoints a single 5G NR cell can bridge before per-flow QoS degrades, showing that sub-3 ms TSN deadlines may require radio-configuration changes such as configured grants or higher numerology.
Mohamed A. M. Seliem, U. Roedig, C. Sreenan et al.· 0 citations
QPS-ToR is proposed, which replaces NegotiaToR's scheduling logic with SW-QPS, a sliding-window algorithm originally proposed for crossbar scheduling that achieves around 90% throughput with a single low-complexity iteration.
The lack of determinism restricts the integration of safety-critical applications into Edge–Fog–Cloud (EFC) architectures. Existing EFC schedulers are typically designed for dynamic, best-effort operation based on unmanaged resource allocation and elastic virtualization. This paradigm introduces unbounded queueing, res...
Omar Hekal, Josepaul Paulachan, Daniel Onwuchekwa et al.· Future Internet· 0 citations
This work presents CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs) that combines spatial and temporal parallelism to perform reductions without accumulation buffers.
Sumukh Pinge, Hardik Soni, Bob Lantz et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.