The results show that QAOA obtains modest but consistent latency reductions, highlighting its value as a diagnostic benchmark for studying the interaction between algorithm structure, entanglement management, and quantum-network architecture.
Abstract
Quantum data-center (QDC) architectures aim to scale distributed quantum computing (DQC) by interconnecting multiple quantum processing units (QPUs), but their performance depends strongly on how algorithmic communication patterns interact with entanglement generation, switch reconfiguration, and network topology. This paper studies the Quantum Approximate Optimization Algorithm (QAOA) as a graph-structured optimization workload for QDC-based distributed quantum computing. We adapt QAOA to SwitchQNet, a distributed quantum compiler framework that schedules communication and entanglement generation over switch-based QDC networks, by adding a routing generator that converts graph-dependent two-qubit cost interactions into remote-CX communication requests across QPUs. Using this extension, we evaluate QAOA instances across Clos, fat-tree, and spine-leaf topologies, measuring communication latency, EPR-pair overhead, EPR wait time, retry overhead, and sensitivity to buffer size, look-ahead depth, communication-qubit count, EPR latency, and EPR fidelity assumptions. The results show that QAOA obtains modest but consistent latency reductions, highlighting its value as a diagnostic benchmark for studying the interaction between algorithm structure, entanglement management, and quantum-network architecture.
Distributed Quantum Computing (DQC) enables scalable quantum execution by interconnecting multiple quantum processing units (QPUs) through quantum networks. In DQC, end-to-end performance is jointly affected by circuit partitioning, entanglement routing, scheduling, and heterogeneous hardware characteristics. However, existing studies often optimize these components independently, providing limited understanding of their cross-layer interactions. In this paper, we present a cross-layer joint-optimization study for DQC using the previously developed SimDisQ-Net simulator. Through simulations, we find that circuit orchestration is one of the dominant factors affecting distributed execution quality, while network-layer mechanisms provide secondary but still meaningful improvements. We further demonstrate that traditional communication metrics, such as hop count or remote-gate count alone, are insufficient predictors of execution quality due to the strong interaction among path fidelity, hardware characteristics, and circuit structure. Motivated by these findings, we propose a topology-aware fidelity proxy (TAFP) evaluation approach that approximates distributed execution fidelity, enabling efficient evaluation of candidate circuit optimizations without time-consuming simulation. Our results highlight the importance of integrated circuit-network orchestration for scalable DQC.
Yeong Lim Tan, Sen Zhang, Haneen Alfauri et al.· Proceedings of the 3rd ACM S...· 0 citations
Distributed Quantum Computing (DQC) addresses the physical scaling limitations of monolithic quantum processors by networking modular Quantum Processing Units (QPUs). Efficient execution of quantum algorithms on DQC architectures requires compiling them across QPUs while minimizing inter-QPU communication bottlenecks, primarily through circuit partitioning. However, current evaluations of state-of-the-art partitioning heuristics focus primarily on the total entanglement cost of the partitions, failing to capture the broader structural and temporal overheads introduced by distributed network constraints. This paper addresses this evaluation gap by applying established monolithic benchmarking metrics to partitioned distributed circuits to quantify the performance impact of network constraints. Using an open-source, automated evaluation pipeline, we systematically assess diverse partitioning algorithms across standardized workloads and quantum network topologies. Our empirical results reveal that partitioning algorithms with comparable entanglement costs can still introduce drastically different physical execution penalties. By exposing these hidden trade-offs, such as severe increases in circuit depth and substantial reductions in gate density, this study demonstrates that comprehensive circuit-level metrics are essential for guiding the future design of DQC compilers.
Javier Vela-Tambo, Davud Azizov, Tian Guo· 0 citations
Distributed quantum computing (DQC) offers a promising approach to scale quantum computing by overcoming the resource limitations of a single quantum processor. However, inter-node communication remains a major bottleneck of DQC due to inefficient and error-prone entanglement distribution. Optimizing inter-node communication can not only reduce the amount of entanglement resource needed to execute a quantum circuit but also improve execution speed and accuracy of the results. This paper proposes DPRQ, a qubit routing algorithm for minimizing inter-node communication in distributed quantum circuits divided into collective communication blocks. Unlike current approaches that utilize greedy block-level qubit routing strategies, DPRQ employs a dynamic programming-based technique focused on global circuit-level optimization, while capturing inter-block dependencies. We evaluated DPRQ on four sets of quantum circuits and a variety of DQC configurations. The results demonstrate that DPRQ's innovative routing strategy achieves an average of 24.40% reduction with a maximum of 85.06% reduction in inter-node communication, when compared to the state-of-the-art collective communication-based DQC compiler QuComm.
Quantum internet applications coordinate classical control, local quantum operations, and entanglement/communication across independent nodes under strict protocol ordering constraints. While Qoala and QNodeOS provide programming and execution abstractions for such applications, there is no compiler infrastructure that systematically lowers Qoala programs to executable node-level code while enabling optimization. We present the first multi-level compiler pipeline for Qoala quantum internet programs. The compiler lowers programs through three MLIR-based IRs that progressively expose hybrid control flow, explicit quantum memory, and Qoala block structure. On top of this pipeline, we implement (i) local peephole rewrites, and (ii) a block-reordering pass formulated as a MILP to reduce qubit lifetimes under precedence and resource constraints. We also provide static analyses for gate counts, qubit lifetimes, fidelity estimation, and quantum-memory efficiency. We evaluate on rotation-merging and blind quantum computing benchmarks using density-matrix simulation and IRderived estimates, showing that block reordering improves fidelity and memory reuse in latency-dominated regimes with modest compilation overhead.
Sacha Bernheim, Bart van der Vecht, Davide Ferrari et al.· 2026 IEEE International Conf...· 0 citations
A comparative study of three entanglement management paradigms for multi-core quantum processors shows that adaptive entanglement managements can substantially improve communication efficiency in scalable quantum multi-core systems.
Rajeswari Suance, Anubhab Dutta, Ruchika Gupta et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.