AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport provides fast amortized samples, but reward tuning can reshape the generated distribution rather than improve exploration of this fixed target. We introduce GRPO-QPS, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training. Across the evaluated reconstruction benchmarks, GRPO-QPS improves over BuresTomFlow and Flow-GRPO on thermal, cat, Dicke, and cluster families, a

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:21:05.240Z. This is not the publication date.

Observed changes

AIIC observation times, not verified publisher revision times. Up to eight recent revisions.

2026-09-24T10:22:25.365Z

  • title: GRPO-QM: Target Preserving Exploration for Quantum Tomography → GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling
  • summary: Reward-based learning can alter the very posterior distribution that scientific inference aims to estimate. GRPO-QM sidesteps this by learning only an exploration strategy for a stated quantum-tomography posterior: a group-relative policy chooses among reversible physical moves, and an exact Metropolis correction ensures the posterior remains stationary once the policy is fixed. We then examine what learning contributes beyond physical proposal mechanisms and prior knowledge. Reconstruction comparisons suggest that most of the gains over the tested flows come from those two components rather t → Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport provides fast amortized samples, but reward tuning can reshape the generated distribution rather than improve exploration of this fixed target. We introduce GRPO-QPS, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training. Across the evaluated reconstruction benchmarks, GRPO-QPS improves over BuresTomFlow and Flow-GRPO on thermal, cat, Dicke, and cluster families, a