The results show that circuit-level parallelism can reduce execution cost on current quantum hardware and simulation time on GPU clusters, with the potential for greater benefits as device quality and qubit counts increase.
Abstract
Today's quantum processors have tens to hundreds of physical qubits, but reliable execution of arbitrary circuits remains limited to fewer than 30 entangled qubits across hardware modalities. Building upon prior work, we introduce an error- and topology-aware method for mapping multiple independent circuits onto disjoint regions of a single large-scale QPU for parallel circuit execution. For applications with many similarly sized circuits, such as observable estimation for Hamiltonian simulation, this approach can reduce billed QPU execution time, with ideal speedup proportional to the number of usable partitions. We demonstrate the approach on IBM's 156-qubit ibm_boston processor using standard QED-C benchmark and Hamiltonian-based observable-estimation workloads. Compared with standard sequential execution, parallel execution reduces billed execution time by 3.5-5.5x while retaining 83-92% of the sequential fidelity. We further evaluate the scaling of parallel circuit execution using GPU-accelerated classical simulation, distributing measurement circuits across GPUs via MPI. Using CUDA-Q on the NERSC Perlmutter system, we achieve up to 13.8x speedup on 16 GPUs (86% parallel efficiency) for an H2 electronic-structure simulation, with scaling evaluated across multiple Hamiltonians and circuit counts. These results provide an indication of the performance ceiling that parallel execution on future quantum hardware may eventually approach. Both execution modes are implemented as a runtime option within the QED-C Application-Oriented Benchmark suite. Together, the results show that circuit-level parallelism can reduce execution cost on current quantum hardware and simulation time on GPU clusters, with the potential for greater benefits as device quality and qubit counts increase.
This work formalizes CLOPS_h (Circuit Layer Operations Per Second) as a holistic speed benchmark defined over layered, hardware-aware circuits, and shares its layer decomposition with scalable layer-fidelity (LF) quality benchmarks, enabling coherent interpretation of speed and quality without conflating the two.
This work demonstrates direct, cross-platform measurement of quantum computational capability using a new benchmark that quantifies the size of the largest computationally relevant quantum circuits that a machine can execute successfully and the speed at which it can execute them, and project the growth of capability a...
Timothy Proctor, Oliver Hart, Oliver Widzowski Maupin et al.· 0 citations
Quantum computing has the potential to accelerate various fields by solving specific problems significantly faster than classical computers. Solving more complex problems generally requires a larger number of qubits. However, current quantum devices are constrained by limited qubit counts and environmental noise. Quant...
Po-Hsuan Huang, Chun-Yen Tai, Chia-Heng Tu et al.· ACM Transactions on Design A...· 0 citations
A low-overhead fidelity-aware scheduling framework for multi-QPU systems based on a Graph Neural Network that estimates, before compilation, the expected fidelity of each circuit on each available QPU, and a tunable scheduler uses these estimates to control the trade-off between execution fidelity and parallelism.
Innocenzo Fulginiti, Antonio Tudisco, Salvatore Zammuto et al.· 0 citations
The realization of practical quantum advantage requires executing large-scale circuits that far exceed the qubit capacity of any single quantum processor. To address this, two primary scaling strategies have emerged: circuit cutting, which utilizes classical resources to decompose circuits into smaller fragments, and m...
Ze-Fan Du, Wen-Rui Zhang, Jake Gesseck et al.· IEEE International Conferenc...· 0 citations
2-local qubit Hamiltonian simulation, a fundamental task in quantum computing, is widely applied in various applications. This paper presents QBX, the first quantum compiler designed for 2-local qubit Hamiltonian simulation on quantum chiplet architectures. Existing general-purpose quantum compilers for chiplet archite...