Skip to content
Conference

Calibration of Variational Quantum Classifiers Under Depolarizing Noise: Expected Calibration Error, Ansatz Expressibility, and Post-Hoc Temperature Scaling

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 297-303 · 0 citations · 17 references

Abstract

Variational Quantum Classifiers (VQCs) have emerged as prime candidates for machine learning on NISQ systems. It stands to reason that the same depolarizing noise that drives quantum states toward the maximally mixed state would also alleviate the overconfidence of VQCs. This paper tests that hypothesis through an empirical study on three datasets at six noise levels, validated across ten random seeds and supported by a formal analysis of how the depolarizing channel contracts the measured logits. The primary finding is that depolarizing noise does not reduce overconfidence in VQCs that remain in the learnable regime: expected calibration error (ECE) never decreases with noise on any dataset, and on the lowest-variance dataset it increases significantly (Wilcoxon signed-rank p < 0.005 over ten seeds). We derive why the optimizer compensates for the channel and confirm the mechanism through a confidence-trajectory experiment. A secondary finding is that a less expressive ansatz can appear well-calibrated only because it collapses to a degenerate solution, demonstrating that ECE must always be reported alongside accuracy. Any effect of noise on accuracy is small and seed-dependent, and it is decoupled from calibration. Post-hoc temperature scaling reduces VQC ECE by 64 to 77 percent across all datasets and is the recommended calibration method for NISQ-era classifiers.

View source

Similar papers

Preprint Jul 2026

Cautious optimism for deep parameterized quantum circuits

A central challenge in quantum machine learning is understanding the scaling behavior of parameterized quantum circuits (PQCs). In particular, it remains unclear how their performance on unseen data changes as the number of trainable parameters increases. Prior works have derived formal generalization guarantees for quantum models, but it is well-known that many such results do not fully characterize generalization behavior in practice. In this work, we show that gradient-based PQCs can exhibit improved performance on unseen data as model size increases, displaying the phenomenon of double descent. This contrasts with the traditional view that larger models lead to degraded generalization. We provide analytical results rigorously underpinning this behavior by leveraging add-one-in perturbation techniques and spectral properties of random matrices. We support these results with numerical experiments on re-uploading PQCs across several data sets and training set sizes, consistently observing the predicted double descent behavior. While other obstacles on the path toward practical quantum machine learning remain, our finding that deeper parameterized quantum circuits do not necessarily exhibit degraded performance provides reasons for cautious optimism.

Marie C. Kempkes, Elies Gil-Fuster, Carlos Bravo-Prieto et al. · 0 citations
Open access Jul 2026

Data quantum Fisher information predicts trainability in variational quantum algorithms

Introduction: Variational quantum algorithms can show unstable trainability and uneven generalisation that are not fully explained by circuit depth or raw parameter count alone. We study the data quantum Fisher information matrix (DQFIM) as a data-dependent effective-capacity diagnostic for controlled supervised unitary-learning tasks. Materials and methods: We used exact-state, classically simulated, 4-qubit, supervised unitary-learning benchmarks with two ansatz families: a hardware-efficient ansatz (HEA) for the baseline trainability study and a symmetry-preserving ansatz for symmetry-controlled comparisons. The analysis covered four linked experiments: E0 trainability phase boundaries, E1 symmetry-controlled data-regime comparisons, E2 trainability versus generalisation, and E3 pre-training prediction of optimisation success. For the main analyses we used the support-basis DQFIM, 30 random seeds per main configuration, grouped cross-validation in the predictive benchmark, and additional diagnostics for threshold sensitivity, phase indeterminacy, Hamming-sector generalisation, and test set size sensitivity. Results: In E0, the empirical trainability boundary increased from M c = 32 at L = 1 to M c = 192 at L = 8 , and the support-basis DQFIM-predicted boundary matched the empirical boundary on the resolved scanned grid under the main analysis setting. In E1, the sector-preserving and sector-broken conditions showed no resolved empirical boundary shift on the scanned grid, with only mild low-L asymmetry in the DQFIM-predicted boundary. In E2, trainability and generalisation separated clearly: some regimes generalised well, while others reached near-zero training loss but retained high test loss. The phase diagnostic supported relative phase indeterminacy as a mechanism for failure on in-sector superpositions after basis-state training. In E3, parameter count alone was a weak predictor of optimisation success, with receiver operating characteristic area under the curve (ROC AUC) = 0.648 ; the structural baseline was stronger, with ROC AUC = 0.958 ; and the DQFIM-enhanced model performed best, with ROC AUC = 0.989 . Conclusions: In these small, idealised, classically simulated matched-family tasks, the support-basis DQFIM provides a useful data-dependent pre-training diagnostic of effective capacity on the retained data support. It tracks resolved trainability boundaries and contributes predictive information beyond raw parameter count and structural metadata. The conclusions remain scoped to controlled small-system simulations; larger-qubit, noisy, hardware-executed, and less-matched settings require further validation before claiming practical scalability.

Shreyosha Ganguly, A. Masta, Shalini Devendrababu et al. · 0 citations
Preprint Jul 2026

Certified Optimal Measurement Reduction over Quantum Context Landscapes

Quantum-measurement reduction contains two distinct global-optimization layers: a continuous problem of splitting an observable and allocating shots within a fixed measurement dictionary, and a nonconvex outer problem of designing the dictionary and calibrating its data-driven uncertainty model. We solve the inner layer globally and certifiably as a second-order cone program (SOCP), and use RANGE, a robust adaptive nature-inspired global optimizer, for the combinatorial and statistical outer layer. For any declared set of contexts, per-shot costs, score functions, and covariance model, the SOCP returns the minimum leading shot cost among unbiased linear stratified estimators. The conic dual supplies an independently checkable lower-bound witness; after feasibility repair, an external verifier recomputes $L \le \Phi \le U$ from stored data without trusting the optimizer. Pilot measurements yield simultaneous finite-sample covariance brackets, and the dual becomes a pricing oracle for omitted contexts. Discrete RANGE searches covering sub-dictionaries, Pareto compression fronts, and candidate contexts; continuous RANGE performs an explicitly empirical, coverage-constrained calibration of covariance-radius models, while rigorous certificates retain the proved finite-sample radius. RANGE compresses molecular context dictionaries by 4.3-6.1x at 0.2-2.1% certified-frontier excess. Standard strategies are exactly optimal for H2 yet leave factors of 2.1-7.7 in shots within their own settings by H2O. Adding fully commuting contexts lowers the certified optimum by up to 56%; on 29-35-qubit production f-element Hamiltonians under a declared Hartree-Fock-proxy covariance model, the capped-dictionary enlargement saves 31-70% of the shots, and transformations reducing block-encoding cost need not reduce sampling cost.

F. Zahariev, Vanda A. Glezakou · 1 citation
Preprint Jul 2026

The Fourier Wall: Why Public Tabular Datasets Refuse Quantum Advantage, and a Certified Recipe for Where It Lives

Across public tabular benchmarks, quantum machine-learning (QML) models usually lose to carefully tuned classical baselines. We argue that this is a structural property of the datasets rather than merely a limitation of current models. Because an angle-encoded quantum neural network is a partial Fourier series, a genuine advantage can arise only when the target spectrum is simultaneously off-grid, of interaction order at least three, high-frequency, supported by near-independent features, and dense beyond practical enumeration. We call the failure to satisfy these conditions the Fourier wall. We operationalize the conditions as SPECTRA, a two-tier certificate: a simulator-free structural screen followed by a decisive comparison between a matched quantum model and five tuned classical twins using paired-bootstrap confidence bounds. On industrial smart-meter data, SPECTRA correctly refuses the real peak-load target, for which gradient-boosted trees reach a held-out ROC-AUC of 0.999. On the same real energy phases with labels generated by an interacting quantum process, the dynamics-matched quantum model reaches 0.994 versus 0.699 for the best generic classical baseline, and collapses to chance when its interaction couplings are ablated. An exact classical simulator ties the quantum model at small width but incurs measured exponential evaluation cost, with an estimated hardware crossover near 13-19 sites. These results provide a practical recipe for identifying, engineering, and deploying quantum-advantage candidates in tabular data.

Javier Mancilla, Tomás Tagliani · 0 citations
Preprint Jul 2026

When cheap gradients fail: the measurement cost of attacking quantum classifiers

Adversarial perturbations threaten machine learning classifiers, including variational quantum classifiers. We show that finite quantum measurement statistics (shot noise) act as a built-in defense against gradient-based test-time attacks whose cost scales unfavorably for the attacker. Because every gradient component must be inferred from repeated circuit executions under any unbiased gradient-estimation rule, white-box extraction consumes a dimension-dependent measurement budget that measurement grouping cannot remove in expressive circuits. Under stated assumptions, single-step attacks need at least quadratically many shots in the input dimension $d$, growing as $d^{5/2}$ under norm-concentration scaling, with a sufficient-budget analysis for iterative attacks via stochastic gradient Langevin dynamics. Simulations up to 784 input dimensions validate the law: the realized total budget is the $d^{5/2}$ geometric floor for plateau-mitigated models and grows as $d^{3.00}$ for the tested deep circuits, whose gradient norms decay with dimension absent barren-plateau mitigation; folding the measured gradient norm back in recovers the parameter-free $d^{3/2}$ shot-noise geometry. Against a matched classical baseline whose attack overhead is dimension-independent (the cheap-gradient principle of automatic differentiation), the quantum gradient cost ratio grows empirically as $d^{3.00}$, so the attacker's relative cost diverges as the model scales. Experiments on a 156-qubit IBM processor (ibm_boston, 4-qubit circuits, $d=12$) reproduce the effect: at matched budgets the device attack tracks the ideal within a few percent, with the high-shot gradient faithful to the exact one. The defense operates precisely when the forward map is classically hard to simulate: only then is a white-box attacker denied the simulate-and-backpropagate shortcut and must pay the measurement cost we quantify.

Bacui Li, Chandra Thapa, Tansu Alpcan et al. · 0 citations