Skip to content

GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling

Sep 2026 · 0 citations · 24 references
Computer Science Physics

TL;DR

GRPO-QPS is introduced, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training, which combines target-preserving Bayesian inference with broad gains over learned transport baselines and a sampling advantage when efficient exploration requires proposal geometry beyond the evaluated conventional kernels.

Abstract

Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport provides fast amortized samples, but reward tuning can reshape the generated distribution rather than improve exploration of this fixed target. We introduce GRPO-QPS, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training. Across the evaluated reconstruction benchmarks, GRPO-QPS improves over BuresTomFlow and Flow-GRPO on thermal, cat, Dicke, and cluster families, and it closely matches an exact two-qubit reference posterior. Tuned conventional MCMC is slightly stronger on several original continuous benchmarks where the available fixed proposals already match the posterior geometry well. To test whether this reflects a fundamental limitation of learned exploration, we evaluate a more challenging multimodal thermal posterior. At six qubits and 800 shots, the learned proposal achieves a minimum effective sample size of 102 per $1{,}000$ likelihood calls, compared with 28 for prior independence, 27 for a tuned fixed mixture, and 20 for Haario adaptive Metropolis. A record-conditioned policy also transfers to unseen 3,000-shot records, matching or exceeding the strongest conventional baseline in all nine held-out seed-record comparisons. These results show that GRPO-QPS combines target-preserving Bayesian inference with broad gains over learned transport baselines and a sampling advantage when efficient exploration requires proposal geometry beyond the evaluated conventional kernels.

View source

Similar papers

#machine learning Preprint Aug 2026

SPSA Hyperparameter Tuning for Variational Quantum Natural Language Inference

Training variational quantum models requires choosing between parameter-shift gradients, which are exact but cost $O(P)$ forward evaluations, and simultaneous perturbation stochastic approximation (SPSA), which uses only two samples but produces high-variance estimates that can degrade optimisation on small supervised...

Nayan D'Souza, Christopher J. Agostino · 0 citations
#artificial intelligence Preprint Sep 2026

Generative Replay Mitigates Sample Starvation in Quantum Architecture Search

GenQAS is introduced, a tensor network-guided RL framework that combines a fixed matrix product state warm-start with prioritized generative replay and can mitigate sample starvation in quantum architecture search and support more resource efficient circuit discovery.

Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld et al. · 0 citations
#machine learning Preprint Sep 2026

Towards Surrogate Based Dequantization of Quantum Reinforcement Learning

In recent years, the utility of parameterized quantum circuits as function approximators has been widely studied. In the context of reinforcement learning, this approach has led to variational quantum algorithms such as quantum Q-learning. While these methods show promising empirical results, and can provide provable a...

Pablo Rodriguez-Grasa, Sofiène Jerbi, Mikel Sanz et al. · 0 citations
Preprint Aug 2026

Automating Variational Quantum Sensing through Reinforcement-Learned Circuit Structures

Numerical results show that the learned architectures recover known benchmark strategies, adapt to dephasing noise, and outperform fixed hardware-efficient ans\"atze while using fewer entangling gates, establishing AutoQSense as a resource-aware approach to adaptive and hardware-compatible quantum sensing.

Jie Liu, Xin Wang · 0 citations
#machine learning Preprint Sep 2026

Quantum MeanFlow: single-shot generative sampling on NISQ hardware

Quantum MeanFlow (QMF), the quantum analogue of the MeanFlow formulation, is established as a viable method for single-step quantum generative sampling, saving on quantum circuit evaluations per generated sample.

Ashish Joshi, Eshaan Mistry, T. Koyama · 0 citations
#machine learning Preprint Sep 2026

Quantum score matching with applications to learning thermal states

This work establishes a general quantum score-matching framework with end-to-end theoretical guarantees for quantum states and achieves information-theoretically optimal sample complexity in the high-temperature regime for Hamiltonians with bounded locality and interaction degree.

Yu-Long Dong, Jia-Qi Leng · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Microsoft Research Blog Aug 20, 2026

Broadening access to Skala creates a faster path to predictive DFT 

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT  appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.