Skip to content
Preprint

Parallel Circuit Execution for Scalable Quantum Computation

Sep 2026 · 0 citations · 52 references
Physics

TL;DR

The results show that circuit-level parallelism can reduce execution cost on current quantum hardware and simulation time on GPU clusters, with the potential for greater benefits as device quality and qubit counts increase.

Abstract

Today's quantum processors have tens to hundreds of physical qubits, but reliable execution of arbitrary circuits remains limited to fewer than 30 entangled qubits across hardware modalities. Building upon prior work, we introduce an error- and topology-aware method for mapping multiple independent circuits onto disjoint regions of a single large-scale QPU for parallel circuit execution. For applications with many similarly sized circuits, such as observable estimation for Hamiltonian simulation, this approach can reduce billed QPU execution time, with ideal speedup proportional to the number of usable partitions. We demonstrate the approach on IBM's 156-qubit ibm_boston processor using standard QED-C benchmark and Hamiltonian-based observable-estimation workloads. Compared with standard sequential execution, parallel execution reduces billed execution time by 3.5-5.5x while retaining 83-92% of the sequential fidelity. We further evaluate the scaling of parallel circuit execution using GPU-accelerated classical simulation, distributing measurement circuits across GPUs via MPI. Using CUDA-Q on the NERSC Perlmutter system, we achieve up to 13.8x speedup on 16 GPUs (86% parallel efficiency) for an H2 electronic-structure simulation, with scaling evaluated across multiple Hamiltonians and circuit counts. These results provide an indication of the performance ceiling that parallel execution on future quantum hardware may eventually approach. Both execution modes are implemented as a runtime option within the QED-C Application-Oriented Benchmark suite. Together, the results show that circuit-level parallelism can reduce execution cost on current quantum hardware and simulation time on GPU clusters, with the potential for greater benefits as device quality and qubit counts increase.

View source

Similar papers

Preprint Aug 2026

CLOPS: Benchmarking System Speed at Utility Scale

This work formalizes CLOPS_h (Circuit Layer Operations Per Second) as a holistic speed benchmark defined over layered, hardware-aware circuits, and shares its layer decomposition with scalable layer-fidelity (LF) quality benchmarks, enabling coherent interpretation of speed and quality without conflating the two.

A. Wack · 0 citations
Preprint Sep 2026

Benchmarking the computational power of quantum computers

This work demonstrates direct, cross-platform measurement of quantum computational capability using a new benchmark that quantifies the size of the largest computationally relevant quantum circuits that a machine can execute successfully and the speed at which it can execute them, and project the growth of capability a...

Timothy Proctor, Oliver Hart, Oliver Widzowski Maupin et al. · 0 citations
Open access Aug 2026

QCutSim: Accelerating Quantum Circuit Cutting Simulation on Consumer-Grade Classical Systems

Quantum computing has the potential to accelerate various fields by solving specific problems significantly faster than classical computers. Solving more complex problems generally requires a larger number of qubits. However, current quantum devices are constrained by limited qubit counts and environmental noise. Quant...

Po-Hsuan Huang, Chun-Yen Tai, Chia-Heng Tu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fidelity-Aware Scheduling of Quantum Circuits on Multi-QPU Systems

A low-overhead fidelity-aware scheduling framework for multi-QPU systems based on a Graph Neural Network that estimates, before compilation, the expected fidelity of each circuit on each available QPU, and a tunable scheduler uses these estimates to control the trade-off between execution fidelity and parallelism.

Innocenzo Fulginiti, Antonio Tudisco, Salvatore Zammuto et al. · 0 citations
Conference Sep 2026

Efficient Circuit Management and Scheduling in Multi-Node Quantum Systems with Dynamic Links

The realization of practical quantum advantage requires executing large-scale circuits that far exceed the qubit capacity of any single quantum processor. To address this, two primary scaling strategies have emerged: circuit cutting, which utilizes classical resources to decompose circuits into smaller fragments, and m...

Ze-Fan Du, Wen-Rui Zhang, Jake Gesseck et al. · 0 citations
Preprint Sep 2026

QBX: A Compiler for 2-local Qubit Hamiltonian Simulation on Quantum Chiplets

2-local qubit Hamiltonian simulation, a fundamental task in quantum computing, is widely applied in various applications. This paper presents QBX, the first quantum compiler designed for 2-local qubit Hamiltonian simulation on quantum chiplet architectures. Existing general-purpose quantum compilers for chiplet archite...

Zi-Kun Li, Zhuo-Ming Chen, Zhi-Hao Jia · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.