REVIEW 1 major objections 53 references
Reinforcement learning designs real-time measurement and feedback for quantum state preparation using only outcome histories.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Reinforcement learning under partial observability with a stochastic terminal reward from one-shot Hamiltonian measurements enables experiment-compatible preparation of Bose-Hubbard ground states and GHZ states.
T0 review reviewed 2026-06-27 challenge →
load-bearing objection The paper frames a POMDP-style RL controller for measurement-feedback state prep with a stochastic single-shot reward that is unbiased by linearity, but the abstract supplies no results at all. the 1 major comments →
Experiment-compatible measurement--feedback quantum state preparation with reinforcement learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
An adaptive measurement-feedback protocol based on reinforcement learning under partial observability prepares ground states of the Bose-Hubbard model and generates GHZ states. The controller uses only the history of experimentally accessible measurement outcomes to choose both the measurement operator and the feedback action in real time. Training uses a stochastic terminal reward built from one-shot measurements of randomly sampled Hamiltonian components that avoids unphysical full-state reconstruction while remaining an unbiased estimator of the target energy.
What carries the argument
Reinforcement learning controller under partial observability that selects measurement operators and feedback actions from measurement history alone, trained via stochastic terminal reward from one-shot Hamiltonian samples.
Load-bearing premise
The stochastic terminal reward built from one-shot measurements of randomly sampled Hamiltonian components is an unbiased estimator of the target energy.
What would settle it
An experiment on the Bose-Hubbard model in which the long-run average of the stochastic reward converges to a value measurably higher than the known ground-state energy of the Hamiltonian.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript develops an adaptive measurement-feedback protocol for quantum state preparation using reinforcement learning under partial observability. The controller selects both the measurement operator and feedback action in real time based solely on the history of experimentally accessible measurement outcomes. Training uses a stochastic terminal reward constructed from one-shot measurements of randomly sampled Hamiltonian components, which is presented as an unbiased estimator of the target energy via linearity of expectation. The approach is demonstrated by preparing ground states of the Bose-Hubbard model and generating GHZ states, positioning it as a scalable, hardware-compatible alternative to methods requiring full-state reconstruction.
Significance. If the numerical demonstrations hold, the work provides a practical advance for experiment-compatible quantum control by eliminating reliance on unphysical full-state information while preserving an unbiased reward signal. The partial-observability POMDP formulation and stochastic reward construction align directly with laboratory constraints, potentially enabling RL-based preparation of correlated and entangled states in many-body systems where tomography is infeasible. This could impact quantum simulation platforms by offering a data-driven, adaptive alternative to handcrafted policies.
major comments (1)
- [Abstract / Results] Abstract and results sections: the central claim of successful demonstrations on Bose-Hubbard ground states and GHZ generation is asserted, yet the manuscript provides no numerical results, training curves, error bars, fidelity metrics, or baseline comparisons. This absence is load-bearing for the claim that the RL controller achieves state preparation and must be remedied with quantitative evidence before the result can be assessed.
Simulated Author's Rebuttal
We thank the referee for their detailed review and for identifying this critical gap in the presentation of our results. We agree that the absence of quantitative evidence undermines the central claims and will revise the manuscript accordingly.
read point-by-point responses
-
Referee: [Abstract / Results] Abstract and results sections: the central claim of successful demonstrations on Bose-Hubbard ground states and GHZ generation is asserted, yet the manuscript provides no numerical results, training curves, error bars, fidelity metrics, or baseline comparisons. This absence is load-bearing for the claim that the RL controller achieves state preparation and must be remedied with quantitative evidence before the result can be assessed.
Authors: We fully agree that the current manuscript lacks the required numerical support for the claimed demonstrations. The abstract and results sections assert successful preparation of Bose-Hubbard ground states and GHZ states without providing training curves, error bars, fidelity values, or baseline comparisons. This is a substantive omission that prevents proper evaluation of the method. In the revised manuscript we will add a dedicated results section containing: (i) learning curves showing reward and fidelity versus training episodes with error bars from multiple random seeds; (ii) final-state fidelities (or energy deviations) for both the Bose-Hubbard and GHZ tasks; (iii) comparisons against at least one baseline policy (e.g., random or hand-crafted feedback); and (iv) representative measurement-outcome histograms confirming that the stochastic reward remains unbiased. These additions will be placed in the main text with appropriate figure captions and will directly substantiate the abstract claims. revision: yes
Circularity Check
No significant circularity
full rationale
The paper applies standard reinforcement learning in a POMDP setting to quantum feedback control. The central technical claim—that a stochastic terminal reward formed from single-shot measurements of randomly sampled Hamiltonian terms is an unbiased estimator of the target energy—follows directly from linearity of expectation and requires no fitted parameters, self-citations, or ansatzes that reduce to the target result. No load-bearing derivation step collapses by construction to its own inputs; the method remains externally falsifiable via standard RL theory and experimental benchmarks.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of Experiment-compatible measurement--feedback quantum state preparation with reinforcement learning." pith.science (2026). https://pith.science/paper/B3VEWCY6
@misc{pith2026260613005,
author = {Pith},
title = {Pith review of: Experiment-compatible measurement--feedback quantum state preparation with reinforcement learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/B3VEWCY6}},
note = {Machine review of arXiv:2606.13005}
}
read the original abstract
Ground-state preparation is a critical task in quantum simulation and quantum computing, as it enables the study of correlated phases and the generation of entangled resource states. While measurement--feedback control has emerged as a promising route to state preparation, existing schemes either rely on handcrafted, task-specific policies or are designed using full quantum-state information that is unavailable in real experiments and becomes impractical for large many-body systems. Here we develop an adaptive measurement--feedback protocol based on reinforcement learning under partial observability. The controller uses only the history of experimentally accessible measurement outcomes to choose both the measurement operator and the feedback action in real time. To make training compatible with experiments, we introduce a stochastic terminal reward built from one-shot measurements of randomly sampled Hamiltonian components, avoiding unphysical full-state reconstruction while remaining an unbiased estimator of the target energy. We demonstrate the method by preparing ground states of the Bose--Hubbard model and by generating GHZ states, establishing a scalable and hardware-compatible route to quantum state preparation.
Figures
Reference graph
Works this paper leans on
-
[1]
Thin col- ored lines show 100 sampled trajectories, the bold line is the ensemble mean, and the shaded region indicates±1 standard deviation over 1000 episodes
regimes, with measurement strengthγ/J= 0.3. Thin col- ored lines show 100 sampled trajectories, the bold line is the ensemble mean, and the shaded region indicates±1 standard deviation over 1000 episodes. The dashed line marks the ex- act ground-state energyE gs. the measured observable and feedback operator to be ˆct = X j αt,j ˆnj, ˆFt = (βt,1 +iβ t,2) ...
-
[2]
Pro- grammable quantum annealing architectures with ising quantum wires.PRX Quantum, 1(2):020311, November 2020
Xingze Qiu, Peter Zoller, and Xiaopeng Li. Pro- grammable quantum annealing architectures with ising quantum wires.PRX Quantum, 1(2):020311, November 2020
2020
-
[3]
Lukin, and Hannes Pichler
Lisa Bombieri, Zhongda Zeng, Roberto Tricarico, Rui Lin, Simone Notarnicola, Madelyn Cain, Mikhail D. Lukin, and Hannes Pichler. Quantum adiabatic opti- mization with rydberg arrays: Localization phenomena and encoding strategies.PRX Quantum, 6(2):020306, April 2025. 5
2025
-
[4]
Universal quantum optimization with cold atoms in an optical cavity.Phys
Meng Ye, Ye Tian, Jian Lin, Yuchen Luo, Jiaqi You, Ji- azhong Hu, Wenjun Zhang, Wenlan Chen, and Xiaopeng Li. Universal quantum optimization with cold atoms in an optical cavity.Phys. Rev. Lett., 131(10):103601, September 2023
2023
-
[5]
He-Ran Wang, Dong Yuan, Shun-Yao Zhang, Zhong Wang, Dong-Ling Deng, and L.-M. Duan. Embedding quantum many-body scars into decoherence-free sub- spaces.Phys. Rev. Lett., 132(15):150401, April 2024
2024
-
[6]
Chivilikhin, A
D. Chivilikhin, A. Samarin, V. Ulyantsev, I. Iorsh, A. Oganov, and O. Kyriienko. Mog-vqe: Multiobjective genetic variational quantum eigensolver.arXiv: Quan- tum Physics, July 2020
2020
-
[7]
Anastasiou, Yanzhu Chen, Nicholas J
Panagiotis G. Anastasiou, Yanzhu Chen, Nicholas J. Mayhall, Edwin Barnes, and Sophia E. Economou. Tetris-adapt-vqe: An adaptive algorithm that yields shallower, denser circuit ans\”atze.Phys. Rev. Res., 6(1):013254, March 2024
2024
-
[8]
Ryabinkin, Scott N
Ilya G. Ryabinkin, Scott N. Genin, and Artur F. Iz- maylov. Constrained variational quantum eigensolver: Quantum computer search engine in the fock space.J. Chem. Theory Comput., 15(1):249–255, January 2019
2019
-
[9]
Parrish, Edward G
Robert M. Parrish, Edward G. Hohenstein, Peter L. McMahon, and Todd J. Mart´ ınez. Quantum computa- tion of electronic transitions using a variational quantum eigensolver.Phys. Rev. Lett., 122(23):230401, June 2019
2019
-
[10]
Ac- celerated variational quantum eigensolver.Phys
Daochen Wang, Oscar Higgott, and Stephen Brierley. Ac- celerated variational quantum eigensolver.Phys. Rev. Lett., 122(14):140504, April 2019
2019
-
[11]
X. Zhou, I. Dotsenko, B. Peaudecerf, T. Rybarczyk, C. Sayrin, S. Gleyzes, J. M. Raimond, M. Brune, and S. Haroche. Field locked to a fock state by quantum feedback with single photon corrections.Phys. Rev. Lett., 108(24):243602, June 2012
2012
-
[12]
Mu˜ noz-Arias, Pablo M
Manuel H. Mu˜ noz-Arias, Pablo M. Poggi, Poul S. Jessen, and Ivan H. Deutsch. Simulating nonlinear dynamics of collective spins via quantum measurement and feedback. Phys. Rev. Lett., 124(11):110503, March 2020
2020
-
[13]
D. A. Ivanov, T. Yu. Ivanova, S. F. Caballero-Benitez, and I. B. Mekhov. Feedback-induced quantum phase transitions using weak measurements.Phys. Rev. Lett., 124(1):010603, January 2020
2020
-
[14]
Greve, Baochen Wu, James K
Athreya Shankar, Graham P. Greve, Baochen Wu, James K. Thompson, and Murray Holland. Continuous real-time tracking of a quantum phase below the stan- dard quantum limit.Phys. Rev. Lett., 122(23):233602, June 2019
2019
-
[15]
Quantum feedback: Theory, experiments, and applications.Physics Reports, 679:1–60, March 2017
Jing Zhang, Yu-xi Liu, Re-Bing Wu, Kurt Jacobs, and Franco Nori. Quantum feedback: Theory, experiments, and applications.Physics Reports, 679:1–60, March 2017
2017
-
[16]
H. M. Wiseman. Quantum theory of continuous feedback. Phys. Rev. A, 49(3):2133–2150, March 1994
1994
-
[17]
Sudhir, D
V. Sudhir, D. J. Wilson, R. Schilling, H. Sch¨ utz, S. A. Fedorov, A. H. Ghadimi, A. Nunnenkamp, and T. J. Kip- penberg. Appearance and disappearance of quantum cor- relations in measurement-based feedback control of a me- chanical oscillator.Phys. Rev. X, 7(1):011001, January 2017
2017
-
[18]
Nishimori’s cat: Stable long-range entanglement from finite-depth unitaries and weak measurements.Phys
Guo-Yi Zhu, Nathanan Tantivasadakarn, Ashvin Vish- wanath, Simon Trebst, and Ruben Verresen. Nishimori’s cat: Stable long-range entanglement from finite-depth unitaries and weak measurements.Phys. Rev. Lett., 131(20):200201, November 2023
2023
-
[19]
Hierarchy of topological order from finite-depth unitaries, measurement, and feedforward
Nathanan Tantivasadakarn, Ashvin Vishwanath, and Ruben Verresen. Hierarchy of topological order from finite-depth unitaries, measurement, and feedforward. PRX Quantum, 4(2):020339, June 2023
2023
-
[20]
Cox, Graham P
Kevin C. Cox, Graham P. Greve, Joshua M. Weiner, and James K. Thompson. Deterministic squeezed states with collective measurements and feedback.Phys. Rev. Lett., 116(9):093602, March 2016
2016
-
[21]
Andrew C. J. Wade, Jacob F. Sherson, and Klaus Mølmer. Squeezing and entanglement of density oscil- lations in a bose-einstein condensate.Phys. Rev. Lett., 115(6):060401, August 2015
2015
-
[22]
Unconditional quantum-noise suppression via measurement-based quantum feedback.Phys
Ryotaro Inoue, Shin-Ichi-Ro Tanaka, Ryo Namiki, Takahiro Sagawa, and Yoshiro Takahashi. Unconditional quantum-noise suppression via measurement-based quantum feedback.Phys. Rev. Lett., 110(16):163602, April 2013
2013
-
[23]
Rist` e, M
D. Rist` e, M. Dukalski, C. A. Watson, G. de Lange, M. J. Tiggelman, Ya M. Blanter, K. W. Lehnert, R. N. Schouten, and L. DiCarlo. Deterministic entanglement of superconducting qubits by parity measurement and feedback.Nature, 502(7471):350–354, October 2013
2013
-
[24]
Real- time quantum feedback prepares and stabilizes photon number states.Nature, 477(7362):73–77, September 2011
Cl´ ement Sayrin, Igor Dotsenko, Xingxing Zhou, Bruno Peaudecerf, Th´ eo Rybarczyk, S´ ebastien Gleyzes, Pierre Rouchon, Mazyar Mirrahimi, Hadis Amini, Michel Brune, Jean-Michel Raimond, and Serge Haroche. Real- time quantum feedback prepares and stabilizes photon number states.Nature, 477(7362):73–77, September 2011
2011
-
[25]
Poulsen, and Klaus Mølmer
Antonio Negretti, Uffe V. Poulsen, and Klaus Mølmer. Quantum superposition state production by contin- uous observations and feedback.Phys. Rev. Lett., 99(22):223601, November 2007
2007
-
[26]
Quantum feedback control for deterministic entangled photon generation.Phys
Masahiro Yanagisawa. Quantum feedback control for deterministic entangled photon generation.Phys. Rev. Lett., 97(19):190201, November 2006
2006
-
[27]
Bounding fidelity in quantum feedback control: Theory and applications to dicke state preparation.Quantum Sci
Eoin O’Connor, Hailan Ma, and Marco G Genoni. Bounding fidelity in quantum feedback control: Theory and applications to dicke state preparation.Quantum Sci. Technol., 10(3):035049, June 2025
2025
-
[28]
Preparing quantum states by measurement-feedback control with bayesian optimization.Front
Yadong Wu, Juan Yao, and Pengfei Zhang. Preparing quantum states by measurement-feedback control with bayesian optimization.Front. Phys., 18(6):61301, July 2023
2023
-
[29]
Reinforcement learning for quantum technology, January 2026
Marin Bukov and Florian Marquardt. Reinforcement learning for quantum technology, January 2026
2026
-
[30]
Quantum circuit discovery for fault-tolerant logical state preparation with reinforcement learning.Phys
Remmy Zen, Jan Olle, Luis Colmenarez, Matteo Puviani, Markus M¨ uller, and Florian Marquardt. Quantum circuit discovery for fault-tolerant logical state preparation with reinforcement learning.Phys. Rev. X, 15(4):041012, Oc- tober 2025
2025
-
[31]
Koch, Ugo Boscain, Tommaso Calarco, Gunther Dirr, Stefan Filipp, Steffen J
Christiane P. Koch, Ugo Boscain, Tommaso Calarco, Gunther Dirr, Stefan Filipp, Steffen J. Glaser, Ron- nie Kosloff, Simone Montangero, Thomas Schulte- Herbr¨ uggen, Dominique Sugny, and Frank K. Wil- helm. Quantum optimal control in quantum technolo- gies. strategic report on current status, visions and goals for research in europe.EPJ Quantum Technol., 9...
2022
-
[32]
Controlling nonergodicity in quantum many-body systems by reinforcement learn- ing
Li-Li Ye and Ying-Cheng Lai. Controlling nonergodicity in quantum many-body systems by reinforcement learn- ing. https://arxiv.org/abs/2408.11989v3, August 2024
-
[33]
Manipulation of Spin Dynamics by Deep Reinforcement Learning Agent
Jun-Jie Chen and Ming Xue. Manipulation of spin dynamics by deep reinforcement learning agent. https://arxiv.org/abs/1901.08748v2, January 2019. 6
work page internal anchor Pith review Pith/arXiv arXiv 1901
-
[34]
Tak- ing gradients through experiments: Lstms and mem- ory proximal policy optimization for black-box quantum control
Moritz August and Jos´ e Miguel Hern´ andez-Lobato. Tak- ing gradients through experiments: Lstms and mem- ory proximal policy optimization for black-box quantum control. In Rio Yokota, Mich` ele Weiland, John Shalf, and Sadaf Alam, editors,High Performance Computing, pages 591–613, Cham, 2018. Springer International Pub- lishing
2018
- [35]
-
[36]
Clas- sifying global state preparation via deep reinforcement learning.Mach
Tobias Haug, Wai-Keong Mok, Jia-Bin You, Wenzu Zhang, Ching Eng Png, and Leong-Chuan Kwek. Clas- sifying global state preparation via deep reinforcement learning.Mach. Learn.: Sci. Technol., 2(1):01LT02, De- cember 2020
2020
-
[37]
Reinforcement learning for quantum control under physical constraints
Jan Ole Ernst, Aniket Chatterjee, Tim Franzmeyer, and Axel Kuhn. Reinforcement learning for quantum control under physical constraints. https://arxiv.org/abs/2501.14372v2, January 2025
-
[38]
Faster state preparation across quantum phase transition assisted by reinforcement learning.Phys
Shuai-Feng Guo, Feng Chen, Qi Liu, Ming Xue, Jun-Jie Chen, Jia-Hao Cao, Tian-Wei Mao, Meng Khoon Tey, and Li You. Faster state preparation across quantum phase transition assisted by reinforcement learning.Phys. Rev. Lett., 126(6):060401, February 2021
2021
-
[39]
A tutorial on optimal control and reinforcement learning methods for quantum tech- nologies.Physics Letters A, 434:128054, May 2022
Luigi Giannelli, Sofia Sgroi, Jonathon Brown, Gheo- rghe Sorin Paraoanu, Mauro Paternostro, Elisabetta Pal- adino, and Giuseppe Falci. A tutorial on optimal control and reinforcement learning methods for quantum tech- nologies.Physics Letters A, 434:128054, May 2022
2022
-
[40]
Duncan, Pablo M
Callum W. Duncan, Pablo M. Poggi, Marin Bukov, Nikolaj Thomas Zinner, and Steve Campbell. Tam- ing quantum systems: A tutorial for using shortcuts- to-adiabaticity, quantum optimal control, and reinforce- ment learning.PRX Quantum, 6(4):040201, October 2025
2025
-
[41]
Reinforcement learning optimization of the charging of a dicke quantum battery
Paolo Andrea Erdman, Gian Marcello Andolina, Vitto- rio Giovannetti, and Frank No´ e. Reinforcement learning optimization of the charging of a dicke quantum battery. Phys. Rev. Lett., 133(24):243602, December 2024
2024
-
[42]
Artificially intelligent maxwell’s demon for optimal control of open quantum systems.Quantum Sci
Paolo A Erdman, Robert Czupryniak, Bibek Bhandari, Andrew N Jordan, Frank No´ e, Jens Eisert, and Giacomo Guarnieri. Artificially intelligent maxwell’s demon for optimal control of open quantum systems.Quantum Sci. Technol., 10(2):025047, March 2025
2025
-
[43]
Deep reinforcement learning for quantum state preparation with weak nonlinear measure- ments.Quantum, 6:747, June 2022
Riccardo Porotti, Antoine Essig, Benjamin Huard, and Florian Marquardt. Deep reinforcement learning for quantum state preparation with weak nonlinear measure- ments.Quantum, 6:747, June 2022
2022
-
[44]
Machine learning for ground state preparation via measurement and feed- back, February 2025
Chuanxin Wang and Yi-Zhuang You. Machine learning for ground state preparation via measurement and feed- back, February 2025
2025
-
[45]
Reinforcement learning for autonomous preparation of floquet-engineered states: Inverting the quantum kapitza oscillator.Phys
Marin Bukov. Reinforcement learning for autonomous preparation of floquet-engineered states: Inverting the quantum kapitza oscillator.Phys. Rev. B, 98(22):224305, December 2018
2018
-
[46]
V. V. Sivak, A. Eickbusch, B. Royer, S. Singh, I. Tsiout- sios, S. Ganjam, A. Miano, B. L. Brock, A. Z. Ding, L. Frunzio, S. M. Girvin, R. J. Schoelkopf, and M. H. De- voret. Real-time quantum error correction beyond break- even.Nature, 616(7955):50–55, April 2023
2023
-
[47]
V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, and M. H. Devoret. Model-free quantum control with re- inforcement learning.Phys. Rev. X, 12(1):011059, March 2022
2022
-
[48]
Norris, Ants Remm, Michael Kerschbaum, Jean- Claude Besse, Florian Marquardt, Andreas Wallraff, and Christopher Eichler
Kevin Reuer, Jonas Landgraf, Thomas F¨ osel, James O’Sullivan, Liberto Beltr´ an, Abdulkadir Akin, Gra- ham J. Norris, Ants Remm, Michael Kerschbaum, Jean- Claude Besse, Florian Marquardt, Andreas Wallraff, and Christopher Eichler. Realizing a deep reinforcement learning agent for real-time quantum feedback.Nat Com- mun, 14(1):7138, November 2023
2023
-
[49]
Kurt Jacobs and Daniel A. Steck. A straightforward in- troduction to continuous quantum measurement.Con- temporary Physics, 47(5):279–303, September 2006
2006
-
[50]
Discovered policy optimisation
Chris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz, Christian Schroeder de Witt, and Jakob Foerster. Discovered policy optimisation
-
[51]
T. J. Elliott, W. Kozlowski, S. F. Caballero-Benitez, and I. B. Mekhov. Multipartite entangled spatial modes of ultracold atoms generated and controlled by quantum measurement.Phys. Rev. Lett., 114(11):113604, March 2015
2015
-
[52]
Diffraction-unlimited position measurement of ultracold atoms in an optical lattice.Phys
Yuto Ashida and Masahito Ueda. Diffraction-unlimited position measurement of ultracold atoms in an optical lattice.Phys. Rev. Lett., 115(9):095301, August 2015
2015
-
[53]
Cold atoms in cavity- generated dynamical optical potentials.Rev
Helmut Ritsch, Peter Domokos, Ferdinand Bren- necke, and Tilman Esslinger. Cold atoms in cavity- generated dynamical optical potentials.Rev. Mod. Phys., 85(2):553–601, April 2013
2013
This paper was first reviewed by grok-4.3 on June 27, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.