Skip to content

M-CSN: Joint Architecture and Flow Scheduling for Metro-Scale AI Fabric Based on Supernodes

2026 · IEEE Transactions on Network Science and Engineering · Vol 13, pp. 11513-11531 · 0 citations · 78 references

Abstract

Deploying trillion-parameter large language models across metropolitan environments is required to sustain real-time inference. Urban power constraints, however, prohibit monolithic GPU clusters, forcing the integration of distributed supernodes into a citywide compute fabric. Over 100-km distances, optical propagation delays invalidate reactive congestion control for 12.8 Tbps cross-domain pipeline parallelism (PP) traffic. The resulting high bandwidth-delay product causes in-flight data to saturate edge switch buffers before the feedback loop triggers source throttling. We address this limitation with M-CSN, an AI compute architecture featuring LOTUS, a model-driven scheduling engine for metropolitan computing power networks (CPNs). LOTUS replaces delayed reactive signaling with a proactive bi-level scheduling strategy. At the macro-layer, an inverse-SLA mechanism partitions capacity among concurrent flows to guarantee performance isolation. At the micro-layer, the scheduler decouples transmission rates from dynamic window probing: it injects calculated pacing rates and static hyper-BDP windows into source RDMA queue pairs (QPs). This source-side enforcement eliminates feedback lag. Packet-level simulations demonstrate that LOTUS sustains a 92% operational load with PFC-free execution. Eliminating pause-induced jitter and buffer saturation reduces critical-flow latency by 2.6× relative to reactive baselines. This enables fragmented urban resources to operate as a unified, deterministic, supernode-based computing power network.

View source

Similar papers

Book Open access Aug 2026

CSIG: Congestion Signaling for Datacenter Transports

This work introduces CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header, and proposes Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production.

Abhiram Ravi, Nandita Dukkipati, Weiwu Pang et al. · 0 citations
Preprint Aug 2026

Scaling 5G-TSN Bridges: Operating Regimes, Scheduling, and Time Synchronisation Under Heterogeneous Industrial Traffic

The nascTime framework on OMNeT++/Simu5G is used to evaluate how many TSN endpoints a single 5G NR cell can bridge before per-flow QoS degrades, showing that sub-3 ms TSN deadlines may require radio-configuration changes such as configured grants or higher numerology.

Mohamed A. M. Seliem, U. Roedig, C. Sreenan et al. · 0 citations
Open access Aug 2026

Hierarchical Scheduler with Adaptive Time-Budget Reallocation for Time-Triggered Edge-Fog-Cloud Architectures

The lack of determinism restricts the integration of safety-critical applications into Edge–Fog–Cloud (EFC) architectures. Existing EFC schedulers are typically designed for dynamic, best-effort operation based on unmanaged resource allocation and elastic virtualization. This paradigm introduces unbounded queueing, res...

Omar Hekal, Josepaul Paulachan, Daniel Onwuchekwa et al. · 0 citations
Book Open access Aug 2026

CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication

This work presents CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs) that combines spatial and temporal parallelism to perform reductions without accumulation buffers.

Sumukh Pinge, Hardik Soni, Bob Lantz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.