Skip to content
Open access

The role of data-induced randomness in quantum machine learning classification tasks

Jul 2026 · npj Quantum Information · 0 citations

TL;DR

A metric for binary classification tasks, the class margin, is introduced, which analytically connects data-induced randomness with classification accuracy for a given data-embedding map, demonstrating its ability to identify data-induced randomness that hinders classification performance.

Abstract

Quantum machine learning (QML) has surged as a prominent area of research with the objective to go beyond the capabilities of classical machine learning models. A critical aspect of any learning task is the process of data embedding, which directly impacts model performance. Poorly designed data-embedding strategies can significantly impact the success of a learning task. Despite its importance, rigorous analyses of data-embedding effects are limited, leaving many cases without effective assessment methods. In this work, we introduce a metric for binary classification tasks, the class margin , by merging the concepts of average randomness and classification margin. This metric analytically connects data-induced randomness with classification accuracy for a given data-embedding map. We benchmark a range of data-embedding strategies through class margin , demonstrating its ability to identify data-induced randomness that hinders classification performance. We expect this work to provide a new approach to evaluate QML models by their data-embedding processes, addressing gaps left by existing analytical tools

Read PDF

Similar papers

Preprint Jul 2026

Quantum-Enhanced Synthetic Data Generation Using Quantum Circuit Born Machines for Imbalanced Tabular Learning

Findings establish QCBM as a viable complementary tool for data augmentation, particularly for low-dimensional structured tabular data with class imbalance, particularly for low-dimensional structured tabular data with class imbalance.

Tanapol Nuatho, Narisorn Sangnakara, Prapong Prechaprapranwong et al. · 0 citations
Open access Sep 2026

Q-SCOPE: Towards characterizing quantum state geometry for out-of-distribution prediction and benchmark evaluation

Hybrid classical-quantum neural network (HCQNN) models have recently emerged as a powerful approach for classification problems, due to their strong computational representation capabilities coming along with the integration of neural networks and quantum computing. However, these models are not free from the fundamental issue of poor separation between the in-distribution (ID) and Out-Of-Distribution (OOD) samples in the classification output space, which is observed in classical neural network classifiers too. Despite this, to the best of our knowledge, there is no current work that systematically studies the OOD detection problem in the context of the HCQNN models. Towards this end, we benchmark the existing approaches for OOD detection in the classical neural network literature domain on the HCQNN classifier models using the standard datasets and metrics and note their limitations. Thereby we propose a novel strategy suitable for OOD prediction in the HCQNN classifier models using the representational properties of the quantum features’ space. Particularly, we find subspaces within the feature space based on their categorical label information, and make use of a metric called the fidelity score that is extensively used in the quantum computing literature for measuring the similarity between the quantum states. Finally, we substantiate our claims on the efficacy of the fidelity score for OOD detection by demonstrating empirically that we can separate the ID from OOD samples across many standard benchmarks using the class-wise boundary defined characterized by the fidelity scores.

Unknown authors · 0 citations
Preprint Aug 2026

Benchmarking Quantum Feature Encoding Strategies for Binary Classification with QSVM

The way in which classical data are encoded into quantum states plays a significant role in both classification performance and quantum circuit complexity in Quantum Machine Learning. In this study, the effects of different quantum feature encoding strategies on Quantum Support Vector Machine performance were investigated using five binary classification datasets. In particular, the statistical relationships between features were incorporated into quantum circuits through \(RY(\theta)\) and controlled-\(RY(\theta)\) gates, and this approach was compared with conventional quantum feature maps. The results demonstrate that incorporating statistical relationships into the encoding process can influence classification performance. However, more complex and densely entangled circuits do not necessarily yield higher performance. In addition, a composite evaluation metric was employed to jointly assess predictive performance, generalization, and circuit cost. The findings across the five datasets indicate that the choice of quantum feature encoding strategy should account for the underlying structure of the data and that predictive performance should be evaluated together with quantum circuit complexity.

Murat Kurt · 0 citations
Preprint Sep 2026

Discretization-Aware Fine-Tuning for Quantum Machine Learning with Chemical Foundation Models

A key challenge in practical quantum machine learning (QML), particularly for discriminative tasks such as classification, is the limited capacity of near-term quantum devices to encode high-dimensional classical data into small quantum registers. In optimized basis-encoded (bit-bit) settings, this constraint leads to cross-class collisions, where samples with different labels are mapped to the same discrete bit-string and thus become indistinguishable to any downstream model. In this work, we investigate how data representation affects QML performance under such severe information bottlenecks. We introduce discretization-aware fine-tuning (DAFT), a method that adapts a pre-trained chemical foundation model to produce representations that remain informative after quantization. DAFT reduces collision probability through a differentiable soft collision loss. We evaluate both quantum and classical models under a controlled setting in which they receive identical discretized bit-string inputs, isolating the effect of representation from model architecture. On the blood-brain barrier penetration (BBBP) molecular property prediction benchmark using ChemBERTa-77M, DAFT reduces collision counts by several orders of magnitude and improves quantum classification accuracy by more than 12 percentage points compared to a frozen backbone. Importantly, without DAFT, classical models outperform QML under the same input constraints. With DAFT, however, this comparison reverses at higher qubit counts. At 10 qubits, the quantum model surpasses a matched classical baseline trained on identical bit-strings (0.883 vs. 0.855, $p = 0.026$). These results show that, in information-constrained regimes, achieving a quantum advantage critically depends on aligning continuous representations with discrete quantum encodings.

Shunji Matsuura, Sonika Johri · 0 citations
Review Jul 2026

Inherent interpretability provides inherent value in quantum machine learning

The field of quantum machine learning (QML) evolved to value models believed to most directly rival those providing utility in classical ML, namely large-scale neural networks. Although more recently, classical ML has been learning a hard lesson with respect to deploying un-interpretable neural networks in the wild: model interpretability matters for domain-adapted co-design and human adoption. We adopt this larger ML perspective to argue that quantum ML model value can be found through the characterization of its inherent interpretability offerings -- i.e. its mathematical structure that contributes meaningfully to desired model behavior for the specific ML task. To support our perspective, we provide a motivating example of a characterization with quantum Fourier models and random Fourier features (RFF) as approaches to approximate Gaussian process (GP) kernels for uncertainty quantification tasks in ML. The top-down and bottom-up complementarity of the two mathematical constructions reveals that quantum Fourier models offer different tools than RFFs for principled GP kernel design and interpretable discovery for uncertainty quantification with real-world data. To showcase the rich variety of inductive biases enabled by quantum information tools, we review examples from the QML literature -- including symmetry, metric geometry, and topology -- that can be used to design inherently interpretable ML models for specific tasks. We hope this framing encourages the QML community to value the inherent components and mechanisms of quantum models separately from task performance, as inherent interpretability might be the reason that a quantum model, and potentially a quantum computer, gets used in practice for ML.

Kaitlin Gili, Zachary P. Bradshaw · 2 citations
Jul 2026

Cautious optimism for deep parameterized quantum circuits

It is shown that gradient-based PQCs can exhibit improved performance on unseen data as model size increases, displaying the phenomenon of double descent, which contrasts with the traditional view that larger models lead to degraded generalization.

Marie C. Kempkes, Elies Gil-Fuster, Carlos Bravo-Prieto et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.