Skip to content
Book Open access

Dynamic Compute and Network Orchestration for Disaggregated RL

Aug 2026 · Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · 0 citations · 53 references
Computer Science

TL;DR

This work builds Silverstone to orchestrate dynamically both compute and network in disaggregated RL, using a reconfigurable optical-electrical fabric called RFabric that achieves superior performance-cost efficiency at scale over static Fat-Tree networks.

Abstract

Disaggregating the generation and training stages in RL is widely adopted to scale LLM post-training. There are two critical challenges here. First, the generation stage often becomes a bottleneck due to dynamic workload shifts and severe execution imbalances. Second, the decoupled stages result in diverse and dynamic network traffic patterns that strain the conventional static fabric. We build Silverstone to orchestrate dynamically both compute and network in disaggregated RL. Silverstone employs an adaptive compute scheduler that adjusts parallelism configuration to match changing workload characteristics within and across generation steps. Silverstone adopts a reconfigurable optical-electrical fabric called RFabric: It leverages optical circuit switches to reconfigure the aggregation and core layers of the topology on demand, tailoring bandwidth resources to the unique communication patterns across various phases of training, generation, and weight synchronization. Evaluated on a 64-H800 GPU testbed, Silverstone demonstrates up to a 1.42× throughput improvement over static baselines. Using a high-fidelity simulator, we also show that RFabric achieves superior performance-cost efficiency at scale over static Fat-Tree networks.

Read PDF

Similar papers

Book Open access May 2026

Revisiting Bruck: Phase-Efficient All-to-All Collective Communication in Reconfigurable Networks

ReTri, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm is presented, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm that improves completion time and improves reconfigurable Bruck by up to 2.1×.

Anton Juerss, Stefan Schmid · 0 citations
Preprint Jul 2026

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training

BiDiRL, a hybrid time-space multiplexing architecture for asynchronous, disaggregated RL designed to reduce resource idleness, is presented, including a hot-switch runtime that enables rapid switching between rollout and training resources with negligible overhead and a static, scheduling-aware planner based on time-performance modeling.

Zhiqiang Tan, Maoxin Wang, Sijie Wang et al. · 1 citation
Preprint Aug 2026

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

HYDRA jointly explores chiplet composition, placement, inter-chiplet bandwidth provisioning, dynamic batching, dynamic batching, and runtime scheduling, and a fast Markov-based performance estimator that captures multi-tenant runtime dynamics for efficient and accurate exploration.

Jiahao Lin, Alish Kanani, Sang-Won Lee et al. · 0 citations
Book Open access Aug 2026

DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI Clouds

DistDPU is presented, a disaggregated DPU architecture that redefines the scaling abstraction for high-bandwidth cloud networking and co-designs the EM-OM functions and the inter-module fabric to minimize virtualization overhead while enforcing security and manageability invariants equivalent to those of a monolithic DPU.

Hao Mei, Lizhou Gao, Yuanyi Zhu et al. · 0 citations
Book Open access Aug 2026

GeoOrchestra: Orchestrating Heterogeneous Geo-Distributed Training with Network-Aware Scheduling

GeoOrchestra is a system that decouples resource filtering from fine-grained strategy search by abstracting compute nodes via computation and memory profiles while modeling WAN links as a virtual hard pipe, which employs hetero-aware pruning to filter invalid resource sets and a resource-driven search that exploits resource disparities to maximize efficiency.

Ting Liu, Qinghua Wu, Jun Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.