Skip to content
Preprint

Towards Continuous Profiling and Optimization of Quantum-Classical Pipelines

Sep 2026 · 0 citations · 64 references
Physics Computer Science

TL;DR

LLQM (Low-Level Quantum Machine), a profiling-driven meta-framework for quantum-classical pipelines that decomposes pipelines into fine-grained tasks and continuously profiles their CPU/GPU, memory, QPU, and queue dependencies alongside real-time hardware states, is presented.

Abstract

Quantum applications increasingly execute as multi-stage quantum-classical pipelines, interleaving QPU computation with classical stages like circuit generation, transpilation, layout mapping, quantum error mitigation (QEM), and post-processing. These stages have diverse resource requirements and exhibit stochastic behavior under drifting hardware noises, yet existing workflow frameworks treat them as static, isolated components. We present LLQM (Low-Level Quantum Machine), a profiling-driven meta-framework for quantum-classical pipelines. LLQM decomposes pipelines into fine-grained tasks and continuously profiles their CPU/GPU, memory, QPU, and queue dependencies alongside real-time hardware states. This unified runtime abstraction captures cross-stage resource dependencies and reveals how classical and quantum decisions interact, enabling characterization of their impact on fidelity and resource consumption. We evaluate LLQM using QEM as a representative pipeline stage, on IBM 156-qubit Heron r2 processors with circuits up to 100 qubits and 1e7 transpiled gates. Our results show that continuous profiling exposes runtime bottlenecks and enables hardware-, fidelity-, and workload-aware optimizations.

View source

Similar papers

Preprint Aug 2026

CLOPS: Benchmarking System Speed at Utility Scale

This work formalizes CLOPS_h (Circuit Layer Operations Per Second) as a holistic speed benchmark defined over layered, hardware-aware circuits, and shares its layer decomposition with scalable layer-fidelity (LF) quality benchmarks, enabling coherent interpretation of speed and quality without conflating the two.

A. Wack · 0 citations
Preprint Aug 2026

ChainForge: Characterizing Embedding as the Bottleneck in Quantum Annealer Workloads

Quantum Annealers (QAs) are among the first commercially scaled quantum computing systems designed for large-scale optimization. Unlike digital systems that execute sequences of compiled instructions, QAs operate as analog single-instruction machines that directly evolve an Ising Hamiltonian toward low-energy solutions...

Kanishka Jayathilake, Cordelia Brumley, T. Smith et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fidelity-Aware Scheduling of Quantum Circuits on Multi-QPU Systems

A low-overhead fidelity-aware scheduling framework for multi-QPU systems based on a Graph Neural Network that estimates, before compilation, the expected fidelity of each circuit on each available QPU, and a tunable scheduler uses these estimates to control the trade-off between execution fidelity and parallelism.

Innocenzo Fulginiti, Antonio Tudisco, Salvatore Zammuto et al. · 0 citations
Preprint Sep 2026

Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs

Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time q...

Joseph K. L. Lee, M. Malekmohammadi, Hong-Sheng Zheng et al. · 0 citations
Preprint Aug 2026

AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization

The capabilities of AutoQuREO are demonstrated through representative co-design case studies, including early-fault-tolerant quantum algorithms, small error correction codes, gate decomposition and variational training of parametric quantum circuits.

Harshkumar Oza, Aritra Sarkar, S. Abbas et al. · 0 citations
Preprint Aug 2026

QSimAdv: A Late-Bound, Vendor-Agnostic Architecture for High-Performance Quantum-Circuit Simulation

Portability in high-performance quantum-circuit simulation need not begin at the kernel. We present QSimAdv, which makes late binding, rather than a common kernel, the basis of vendor independence. Representation, operator lowering, and data placement are bound only when their required inputs become available. Before f...

Shu-Sen Liu, P. Elahi, Wen Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.