Skip to content
Conference

STON: Scaling Torus-Based AI Training Clusters via Optical Circuit Switches

Jul 2026 · Fall Joint Computer Conference · pp. 169-176 · 0 citations · 30 references

Abstract

Torus networks are deployed in production AI training clusters for their path diversity and low latency, but 2D Torus scales poorly: electrical packet switches compromise latency, and high-dimensional Torus introduces excessive routing complexity. We present STON (Scalable TOrus Network), a hierarchical architecture that treats a 2D Torus as a supernode and interconnects supernodes with a reconfigurable Optical Circuit Switch (OCS) for AlltoAll-dominated large-scale training networks. STON comprises three coordinated modules: (1) fragmentaware task placement, which minimizes inter-supernode traffic by reducing job fragmentation; (2) non-disruptive logical topology mapping, governed by two principles that prevent OCS reconfiguration from disrupting running tasks or partitioning multisupernode jobs; and (3) compute-phase traffic forwarding, which ensures reachability when direct OCS circuits are unavailable. STON reduces average FCT by 42.2%-61.1% across synthetic workloads and by 52.6% on a one-day Kalos production trace (under an AlltoAll traffic model for all jobs), with 95th-percentile tail latency reduced by up to 74.5%, versus a static direct-connect baseline using the same OCS hardware.

View source

Similar papers

Book Open access Aug 2026

CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication

This work presents CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs) that combines spatial and temporal parallelism to perform reductions without accumulation buffers.

Sumukh Pinge, Hardik Soni, Bob Lantz et al. · 0 citations
Book Open access Aug 2026

Dragonfly-Ultra: A Scalable, Low-Cost Network Architecture for High-Performance AI Clusters

Dragonfly-Ultra is presented, a scalable, low-cost network architecture for high-performance AI clusters that can scale to over 260k GPUs with only 82% cost and 81% power consumption of a 3-layer Clos architecture and incorporates three key mechanisms to further improve network performance and optimize collective communication.

Rui Zhuang, Hui Yuan, Junye Zhang et al. · 0 citations
Book Open access May 2026

Revisiting Bruck: Phase-Efficient All-to-All Collective Communication in Reconfigurable Networks

ReTri, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm is presented, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm that improves completion time and improves reconfigurable Bruck by up to 2.1×.

Anton Juerss, Stefan Schmid · 0 citations
2026

TopoCrafter: Toward Highly Reliable Optical-Circuit-Switched HPC/AI Interconnects via Dual-Agent DRL

Optical-circuit-switched interconnects have become one of the core components for AI training due to their flexible topology reconfiguration. In contrast to the applications carried by traditional data center networks, large-scale language model training is highly sensitive to network failures, where frequent disruptions will cause gradient synchronization delays, leading to training interruptions and wasted computational resources. Existing schemes are primarily focused on specific communication patterns, without considering the fault probability distribution. As a result, unreliable links remain on critical paths. Furthermore, passive fault response mechanisms lead to inefficient topology reconfigurations, preventing network protocol convergence and making it difficult to meet the stringent stability requirements of large-scale model training. To address reliability challenges in optical-circuit-switched interconnect, we propose TopoCrafter, which leverages dual-agent deep reinforcement learning to proactively mitigate network failures. The “Topo-Agent” estimates link failure probabilities to determine reconfiguration timing and then employs a lightweight heuristic algorithm to create failure-avoidant topology that matched to traffic pattern. Concurrently, the “Route-Agent” optimizes traffic distribution. Through their strategic interaction, the agents learn holistic policies that optimally balance network reliability and communication efficiency. To improve generalization, a progressive training approach is employed, allowing the agents to adapt to complex failure environments while accelerating convergence. Under link failure scenarios, TopoCrafter maintains reliability, reducing end-to-end latency by up to 50% and maximum link utilization by approximately 20% compared to FatTree. In addition, progressive training algorithm ensures a performance degradation of less than 10% when adapting to new failure environments, and it maintains stable high performance as the network scales.

Liang Qin, Xingyu Liu, Wenting Wei et al. · 0 citations
Book Open access Aug 2026

Balanced Sparse Tree: A Scalable Network Topology for Large Language Models

This work proposes a novel topology named the Balanced Sparse Tree (BST), which is a topology characterized by symmetric design and sparse connections, motivated by hypergraph theory and Steiner Systems, and demonstrates the superiority of BST over the state-of-the-art in network scale, latency, bandwidth, and cost.

Shaoteng Liu, Dejun Kong, Huitian Wang et al. · 0 citations

Abstraction: Flow Prioritization With Spatial Diversity in The Data Center Network

The proposed Multi-Path Multi-Level Feedback Queueing (MP-MLFQ) leverages the spatial diversity and regularity of DCNs to realize a scheduler with numerous logical priority levels while occupying as low as 2 physical priority queues within network switches.

Alessandro Cornacchia, Andrea Bianco, Paolo Giaccone et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.