REVIEW 3 major objections 4 minor 102 references
Reinforcement learning to learn quantum states for Heisenberg scaling accuracy
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A reinforcement-learning agent that tunes an optimizer's step size and learning rate can learn random quantum states with infidelity that falls as one over the total success count, and the three-qubit policy also works on five-qubit states.
desk verdict Useful RL meta-learning for quantum state learning, but the Heisenberg-scaling claim is overstated because the shot accounting omits failure shots and the exponent is largely built into the halting rule. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the evolution-strategy update loop whose two hyperparameters are controlled by the RL agent. At each iteration, the ES samples $k$ parameter perturbations $\vec\theta + \sigma\vec\epsilon_i$ around the current ansatz parameters, measures the success count for each perturbed state, and updates $\vec\theta$ by a reparameterized gradient estimator. The observation fed to the agent is the geometric success count $C$ from the single-shot measurement, which ties the objective directly to measurement statistics rather than to expensive fidelity estimation. The action repetition strategy is the training device that makes RL tractable: repeating a chosen $(\sigma,\eta)$ for $t_{\mathrm{rep}}$ consecutive steps shortens the effective decision horizon early in training, producing a curriculum that lets the actor-critic learn to reduce the halting time.
What would settle it
Record every measurement outcome during a run, set the true shot budget to $S = C_{\mathrm{total}} + (\text{number of measurement sequences})$, and refit $\bar f = \alpha S^{-\beta}$ for 1-, 2-, and 3-qubit states; if the 2- and 3-qubit exponents fall below 1, or the 1-qubit exponent falls clearly below its reported 0.948, the Heisenberg-scaling claim in terms of actual shots does not hold.
Extended reading notes
Core claim
The central claim is that the meta-learned RL policy makes quantum state learning shot-efficient at the statistical limit. The environment is a single-shot measurement scheme: the hardware-efficient ansatz $\hat{U}(\vec\theta)$ is applied to the unknown state, a binary measurement is repeated until a failure outcome occurs, and the number of consecutive successes $C$ is recorded; because $C$ is geometrically distributed with mean $p_s/(1-p_s)$, maximizing $C$ aligns the ansatz with the success basis. The RL agent, an actor-critic pair, observes $C$ and selects the evolution-strategy hyperparameters $\sigma$ (sampling range) and $\eta$ (learning rate), receiving reward $-1$ for each time step before the halting step. With the action repetition strategy, each chosen action is held for $t_{\mathrm{rep}}$ steps, and $t_{\mathrm{rep}}$ is annealed downward during RL training as a curriculum on the depth of the decision process. After training at target success count $C_{\mathrm{target}}=10^4$, the policy reduces average total success counts by 14% (1 qubit), 28% (2 qubit), and 34% (3 qubit) relative to the best fixed action, and the fitted infidelity scaling is $\bar f = \alpha C_{\mathrm{total}}^{-\beta}$ with $\beta \approx 0.948, 1.161, 1.184$; the 3-qubit policy transfers to 4- and 5-qubit states with $\beta \approx -1.189$ and $-0.829$.
Load-bearing premise
The paper equates total success count with total measurement shots, even though each measurement sequence ends in an uncounted failure shot; the scaling claim depends on those failure shots being negligible or scaling at the same rate.
Editorial extensions
If this is right
- If the scaling claim holds, reaching infidelity $\bar f$ costs $O(1/\bar f)$ total success counts, matching the Heisenberg limit and beating the $O(\bar f^{-4/3})$ scaling of standard quantum state tomography.
- At target success count $10^4$, the RL-trained strategy saves 14%, 28%, and 34% of total success counts for 1-, 2-, and 3-qubit states relative to the best fixed hyperparameters, with average fidelities of about 0.99994, 0.99972, and 0.99946.
- The three-qubit-trained policy generalizes to four- and five-qubit Haar-random states, achieving average infidelity around $6.6\times10^{-4}$ and $9.3\times10^{-4}$ with $7.3\times10^6$ and $1.29\times10^8$ success counts respectively.
- The same policy learns a particular entangled five-qubit state to fidelity about 0.9989 with $6.62\times10^7$ success counts, indicating that the learned hyperparameter schedule is not tied to the Haar-random training distribution alone.
- The hardware-efficient ansatz used here has fewer parameters than the general unitary for the same system size, and the ES uses fewer evaluation samples than parameters, so the optimizer-ansatz combination remains practical when gradients are expensive.
Reading between the lines
- Editorial inference: the reported resource is total success count, not total measurement shots; every measurement sequence also ends in a failure shot, so the true shot budget is larger, and refitting exponents against the full shot count is the natural check on the Heisenberg claim.
- Editorial inference: the learned policy seems to implement a simple 'large step size when C is small, small when C is large' schedule; if so, a closed-form adaptive schedule might capture much of the benefit, which could be tested by replacing the RL policy with a hand-coded rule.
- Editorial inference: because the observation is a scalar success count, the same framework could be applied to other success/failure objectives such as gate fidelity or variational eigensolver energies; the open question is whether the scalar observation remains sufficient for those landscapes.
- Editorial inference: the jump from 3-qubit training to 5-qubit states suggests the agent may have learned dimensionless features of the optimization landscape; simulating 6- to 8-qubit states would show whether the scaling exponents stay near 1 or degrade.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a reinforcement-learning-based meta-learning scheme that adapts the hyperparameters (sampling range sigma and learning rate eta) of an evolution strategy used to train a hardware-efficient ansatz for learning unknown pure quantum states via single-shot measurements. The authors report that the trained RL agent reduces the total success count needed to reach a target infidelity compared with fixed-hyperparameter baselines, achieves infidelity scaling close to the Heisenberg limit, and generalizes from 3-qubit training to 4- and 5-qubit state learning. They also introduce an action repetition strategy to make RL training tractable.
Significance. If the scaling claim were fully supported, the work would be a useful demonstration that meta-learned hyperparameter schedules can improve variational quantum state learning and that a hardware-efficient ansatz can be trained with fewer evaluation points than parameters. The manuscript is commendable for releasing code, for presenting the action repetition strategy clearly, and for comparing against maximum-likelihood QST. However, the central resource accounting and the interpretation of the scaling exponents need additional work before the Heisenberg-scaling claim can be accepted.
major comments (3)
- [Algorithm 1, line 15 and Fig. 4] The quantity plotted on the x-axis of Fig. 4 is the total success count C_total, not the total number of measurement shots. In each iteration of Algorithm 1, line 3 measures one sequence and line 8 measures k sequences, and every sequence terminates with a failure outcome. The true shot count is therefore C_total + (k+1)*t_H, where t_H is the number of ES iterations. The paper never reports t_H for the scaling runs. For the training runs at C_target=10^4, the reported t_H values (19, 660, 3409 for 1-, 2-, and 3-qubit states) with k=5, 10, 30 give omitted failure-shot offsets of 114, 7,260, and 105,679, respectively, which are up to about 4.5% of C_total and grow relative to C_total as C_target decreases. For the 5-qubit generalization (k=100, C_total≈1.29e8), t_H is not reported and the omitted failure shots could be comparable to C_total. Since the Heisenberg claim is a fitted exponent on a log-log plot, the authors must report t_H (or the total number of sequences) for every data point and refit the infidelity against C_true = C_total + (k+1)*t_H.
- [§II.A, Eqs. (2)-(4) and §III.C] The near-Heisenberg exponent is largely constrained by the halting rule, so the present data do not demonstrate that the RL agent discovers Heisenberg scaling. Because the infidelity is f = 1 - p_s for a pure state and Algorithm 1 halts when an observed success count satisfies C >= C_target (line 9), the final fidelity is forced to be of order 1/C_target up to geometric fluctuations. The fitted exponent in f = α C_total^{-β} therefore mostly measures how C_total scales with C_target, not an independent scaling law. To support the advertised claim, the authors should report C_total versus C_target for the RL and baseline runs, report the actual final p_s values, and show that the fitted exponent is robust when the x-axis is the true shot count rather than C_total.
- [§III.C, scaling exponents] The reported exponents β ≈ 1.161 and 1.184 for 2- and 3-qubit states exceed the range β ∈ [0.9, 1.0) that the paper itself cites for Heisenberg-limited one-qubit learning, and they also exceed the values one would expect if β=1 were the statistical limit. This inconsistency is likely a consequence of the missing failure shots in the resource accounting, but as written it makes the scaling claim internally questionable. The fits also use only four points per curve with no error bars. The authors should refit with the corrected shot count and provide confidence intervals for β.
minor comments (4)
- [§III.D] The text states that the scaling factor for 4- and 5-qubit generalization is “β = −1.189 and β = −0.829”; negative exponents contradict the fitted form f = α C^{-β} and the plotted decreasing trend. These should be positive values, or the sign convention should be clarified.
- [§II.B and Table I] The action space is defined as A = Aσ × Aη with Aσ, Aη ⊂ [0, ∞), but Table I specifies small discrete sets (e.g., 4 values each). Please state explicitly that the continuous spaces are discretized according to Table I and whether the 169-action case uses the geometric spacing of Eq. (10).
- [Appendix D and Fig. 4] The baseline selection uses “the action which gives the lowest total success count” from a grid of simulations; this is a post-selected baseline. Please state whether the same post-selection was used when computing the improvement percentages in Fig. 4(d), and clarify how simulations that fail within t_max are excluded from the averages.
- [Table I, footnote a] There is a typo in “We obtain similiar results for tl = 200”; it should read “similar.”
Circularity Check
No circularity found: the RL contribution is validated against fixed-action baselines, and the Heisenberg-scaling claim is transparently inherited from the cited SSML protocol.
full rationale
The paper proposes an RL agent that tunes the hyperparameters (sampling range and learning rate) of an evolution strategy used to train a hardware-efficient ansatz for quantum state learning. The reported improvement in shot efficiency is an empirical comparison against fixed-action baselines (Appendix D) chosen as the best grid points without RL; the RL agent is trained to minimize halting time t_H, not to match a pre-fitted target, so this comparison is not circular. The infidelity scaling f ~ O(C^{-1}) is presented as a simulation result obtained by curve-fitting the paper's own data. Although the geometric measurement model (Eq. 2), the reconstruction rule (Eq. 3), and the infidelity definition (Eq. 4) imply that, once a run of length C_target is observed, the success probability must be close to 1 and hence f ~ O(1/C_target), the paper explicitly attributes this single-shot measurement scheme to prior SSML work (Refs. [25,26]) and makes no claim to have derived this scaling from first principles. The actual novelty is the RL-based hyperparameter adaptation and the action repetition strategy, which are evaluated against baselines and shown to reduce total success counts. There is no load-bearing self-citation chain, no uniqueness argument imported from the authors' own prior work, and no fitted parameter relabeled as a prediction. The separate concern that C_total excludes the failure shot terminating each measurement sequence is a correctness and resource-accounting issue, not a circularity, and does not affect the derivation chain itself. Therefore no circular step is identified.
Assumptions & free parameters
free parameters (6)
- ES sample count k =
k = 5, 10, 30, 100 for N = 1, 2, 3, 4, 5
- Action grid A_sigma x A_eta =
16 actions for N=1,2,3; 169 actions for one 2-qubit run
- Action repetition schedule (t_u, t_l, T_th) =
(1,50,100), (80,800,100), (300,2000,500)
- Maximum ES steps t_max used in reward normalization =
3e3, 1e4, 2e4 for 1, 2, 3 qubits
- Power-law scaling fit parameters (alpha, beta) =
beta ~ 0.948, 1.161, 1.184 for 1, 2, 3 qubits; negative sign reported for 4/5 qubits
- HEA depth L =
L = 1, 5, 10 for N = 2, 3, 4/5
assumptions (5)
- standard math Success count C is geometrically distributed with p(C) = p_s^C (1 - p_s).
- domain assumption The reconstructed state fidelity equals p_s.
- domain assumption The ES gradient estimator in Eq. (7) is useful with fewer samples than HEA parameters.
- domain assumption Haar-random pure states are representative target states for the scaling claim.
- domain assumption The HEA with the chosen depth can represent the target random states to the reported infidelity.
Cite this review
Pith. "Pith review of Reinforcement learning to learn quantum states for Heisenberg scaling accuracy." pith.science (2026). https://pith.science/paper/7QUEUUSX
@misc{pith2026241202334,
author = {Pith},
title = {Pith review of: Reinforcement learning to learn quantum states for Heisenberg scaling accuracy},
year = {2026},
howpublished = {\url{https://pith.science/paper/7QUEUUSX}},
note = {Machine review of arXiv:2412.02334}
}
read the original abstract
Learning quantum states is a crucial task for realizing quantum information technology. Recently, neural approaches have emerged as promising methods for learning quantum states. We propose a meta-learning model that utilizes reinforcement learning (RL) to optimize the process of learning quantum states. To improve the data efficiency of the RL, we introduce an action repetition strategy inspired by curriculum learning. The RL agent significantly improves the sample efficiency of learning random quantum states, and achieves infidelity scaling close to the Heisenberg limit. We also show that the RL agent trained using 3-qubit states can generalize to learning up to 5-qubit states. These results highlight the utility of RL-driven meta-learning to enhance the efficiency and generalizability of learning quantum states. Our approach can be applied to improve quantum control, quantum optimization, and quantum machine learning.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
The state st is a quantum state ˆU (⃗θt) |ψ⟩ transformed by the HEA in Fig
S is the Hilbert space of N -qubit states. The state st is a quantum state ˆU (⃗θt) |ψ⟩ transformed by the HEA in Fig. 1 (b)
-
[2]
A is the space of action at = (σt, ηt), where σ and η are hyperprameters of evolution strategy which de- termine sampling range and learning rate, respec- tively [43]
-
[3]
T is the space of transition probability defined by p(st+1|st, at) : S × A × S →[0, 1]
-
[4]
The reward rt = 0 if the quantum state learning is finished at time step t
R is the space of reward rt. The reward rt = 0 if the quantum state learning is finished at time step t. Otherwise, rt = −1
-
[5]
To observe the state, we use a measurement of binary outcome; success and fail
Ω is the space of observation ot. To observe the state, we use a measurement of binary outcome; success and fail. We perform the measurement until a fail outcome appears. The observation is defined as the number of consecutive success outcomes be- fore the fail outcome
-
[6]
O is the space of the transition probability of ob- servation defined by p(ot|st) : O × S →[0, 1]. 11 0.0010.010.11 0.0017.0027.0067.0067.006 0.017.0607.0057.0087.008 0.15.7336.4117.0507.014 17.0927.1097.0847.046 0.0010.010.11 0.001-0.014-0.035-0.038-0.034 0.01-0.302-0.032-0.047-0.052 0.1-3.558-3.620-0.269-0.087 1-1.201-1.205-1.034-0.542 0.0010.010.11 0.0...
-
[7]
Anshu and S
A. Anshu and S. Arunachalam, A survey on the complex- ity of learning quantum states, Nature Reviews Physics 6, 59 (2024)
2024
-
[8]
Gebhart, R
V. Gebhart, R. Santagati, A. A. Gentile, E. M. Gauger, D. Craig, N. Ares, L. Banchi, F. Marquardt, L. Pezz` e, and C. Bonato, Learning quantum systems, Nature Reviews Physics 5, 141 (2023)
2023
Show all 102 references
-
[9]
Vogel and H
K. Vogel and H. Risken, Determination of quasiproba- bility distributions in terms of probability distributions for the rotated quadrature phase, Phys. Rev. A 40, 2847 (1989)
1989
-
[10]
Hradil, Quantum-state estimation, Phys
Z. Hradil, Quantum-state estimation, Phys. Rev. A 55, R1561 (1997)
1997
-
[11]
Paris and J
M. Paris and J. Rehacek, Quantum state estimation , Vol. 13 649 (Springer Science & Business Media, 2004)
2004
-
[12]
Cramer, M
M. Cramer, M. B. Plenio, S. T. Flammia, R. Somma, D. Gross, S. D. Bartlett, O. Landon-Cardinal, D. Poulin, and Y.-K. Liu, Efficient quantum state tomography, Na- ture Communications 1, 149 (2010)
2010
-
[13]
B. P. Lanyon, C. Maier, M. Holz¨ apfel, T. Baumgratz, C. Hempel, P. Jurcevic, I. Dhand, A. S. Buyskikh, A. J. Daley, M. Cramer, M. B. Plenio, R. Blatt, and C. F. Roos, Efficient tomography of a quantum many-body sys- tem, Nature Physics 13, 1158 (2017)
2017
-
[14]
Schwemmer, G
C. Schwemmer, G. T´ oth, A. Niggebaum, T. Moroder, D. Gross, O. G¨ uhne, and H. Weinfurter, Experimental comparison of efficient tomography schemes for a six- qubit state, Phys. Rev. Lett. 113, 040503 (2014)
2014
-
[15]
Aaronson, Shadow tomography of quantum states, in Proceedings of the 50th annual ACM SIGACT sympo- sium on theory of computing (2018) pp
S. Aaronson, Shadow tomography of quantum states, in Proceedings of the 50th annual ACM SIGACT sympo- sium on theory of computing (2018) pp. 325–338
2018
-
[16]
J. M. Lukens, K. J. H. Law, A. Jasra, and P. Lougov- ski, A practical and efficient approach for bayesian quan- tum state estimation, New Journal of Physics 22, 063038 (2020)
2020
-
[17]
Park and M
C.-Y. Park and M. J. Kastoryano, Geometry of learning neural quantum states, Phys. Rev. Res. 2, 023232 (2020)
2020
-
[18]
Ahmed, C
S. Ahmed, C. S´ anchez Mu˜ noz, F. Nori, and A. F. Kockum, Classification and reconstruction of optical quantum states with deep neural networks, Phys. Rev. Res. 3, 033278 (2021)
2021
-
[19]
Ahmed, C
S. Ahmed, C. S´ anchez Mu˜ noz, F. Nori, and A. F. Kockum, Quantum state tomography with conditional generative adversarial networks, Phys. Rev. Lett. 127, 140502 (2021)
2021
-
[20]
P. Cha, P. Ginsparg, F. Wu, J. Carrasquilla, P. L. McMa- hon, and E.-A. Kim, Attention-based quantum tomog- raphy, Machine Learning: Science and Technology 3, 01LT01 (2021)
2021
-
[21]
Lange, M
H. Lange, M. Kebriˇ c, M. Buser, U. Schollw¨ ock, F. Grusdt, and A. Bohrdt, Adaptive Quantum State Tomography with Active Learning, Quantum 7, 1129 (2023)
2023
-
[22]
Gaikwad, O
A. Gaikwad, O. Bihani, Arvind, and K. Dorai, Neural- network-assisted quantum state and process tomography using limited data sets, Phys. Rev. A109, 012402 (2024)
2024
-
[23]
A. M. Palmieri, G. M¨ uller-Rigat, A. K. Srivastava, M. Lewenstein, G. Rajchel-Mieldzio´ c, and M. P lodzie´ n, Enhancing quantum state tomography via resource- efficient attention-based neural networks, Phys. Rev. Res. 6, 033248 (2024)
2024
-
[24]
H. Ma, D. Dong, I. R. Petersen, C.-J. Huang, and G.-Y. Xiang, Neural networks for quantum state tomography with constrained measurements, Quantum Information Processing 23, 317 (2024)
2024
-
[25]
Q.-Q. Wang, S. Dong, X.-W. Li, X.-Y. Xu, C. Wang, S. Han, M.-H. Yung, Y.-J. Han, C.-F. Li, and G.-C. Guo, Efficient learning of mixed-state tomography for photonic quantum walk, Science Advances 10, eadl4871 (2024)
2024
-
[26]
Torlai, G
G. Torlai, G. Mazzola, J. Carrasquilla, M. Troyer, R. Melko, and G. Carleo, Neural-network quantum state tomography, Nature Physics 14, 447 (2018)
2018
-
[27]
Rocchetto, E
A. Rocchetto, E. Grant, S. Strelchuk, G. Carleo, and S. Severini, Learning hard quantum distributions with variational autoencoders, npj Quantum Information 4, 28 (2018)
2018
-
[28]
Carrasquilla, G
J. Carrasquilla, G. Torlai, R. G. Melko, and L. Aolita, Reconstructing quantum states with generative models, Nature Machine Intelligence 1, 155 (2019)
2019
-
[29]
A. M. Palmieri, E. Kovlakov, F. Bianchi, D. Yudin, S. Straupe, J. D. Biamonte, and S. Kulik, Experimen- tal neural network enhanced quantum tomography, npj Quantum Information 6, 20 (2020)
2020
-
[30]
Gao and L.-M
X. Gao and L.-M. Duan, Efficient representation of quan- tum many-body states with deep neural networks, Na- ture Communications 8, 662 (2017)
2017
-
[31]
S. M. Lee, J. Lee, and J. Bang, Learning unknown pure quantum states, Phys. Rev. A 98, 052302 (2018)
2018
-
[32]
S. M. Lee, H. S. Park, J. Lee, J. Kim, and J. Bang, Quan- tum state learning via single-shot measurements, Phys. Rev. Lett. 126, 170504 (2021)
2021
-
[33]
Y. Liu, D. Wang, S. Xue, A. Huang, X. Fu, X. Qiang, P. Xu, H.-L. Huang, M. Deng, C. Guo, X. Yang, and J. Wu, Variational quantum circuits for quantum state tomography, Phys. Rev. A 101, 052316 (2020)
2020
-
[34]
S. Xue, Y. Liu, Y. Wang, P. Zhu, C. Guo, and J. Wu, Variational quantum process tomography of unitaries, Phys. Rev. A 105, 032427 (2022)
2022
-
[35]
Innan, O
N. Innan, O. I. Siddiqui, S. Arora, T. Ghosh, Y. P. Ko¸ cak, D. Paragas, A. A. O. Galib, M. A.-Z. Khan, and M. Ben- nai, Quantum state tomography using quantum machine learning, Quantum Machine Intelligence 6, 28 (2024)
2024
-
[36]
K. H. Wan, O. Dahlsten, H. Kristj´ ansson, R. Gardner, and M. S. Kim, Quantum generalisation of feedforward neural networks, npj Quantum Information 3, 36 (2017)
2017
-
[37]
K. Beer, D. Bondarenko, T. Farrelly, T. J. Osborne, R. Salzmann, D. Scheiermann, and R. Wolf, Training deep quantum neural networks, Nature Communications 11, 808 (2020)
2020
-
[38]
Haghshenas, J
R. Haghshenas, J. Gray, A. C. Potter, and G. K.-L. Chan, Variational power of quantum circuit tensor networks, Phys. Rev. X 12, 011047 (2022)
2022
-
[39]
Abbas, D
A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, The power of quantum neural networks, Nature Computational Science 1, 403 (2021)
2021
-
[40]
Farhi, J
E. Farhi, J. Goldstone, and S. Gutmann, A quan- tum approximate optimization algorithm, arXiv preprint arXiv:1411.4028 (2014)
2014 arXiv
-
[41]
Verdon, M
G. Verdon, M. Broughton, J. R. McClean, K. J. Sung, R. Babbush, Z. Jiang, H. Neven, and M. Mohseni, Learn- ing to learn with quantum neural networks via classical neural networks (2019), arXiv:1907.05415 [quant-ph]
2019 arXiv
-
[42]
Wilson, R
M. Wilson, R. Stromswold, F. Wudarski, S. Hadfield, N. M. Tubman, and E. G. Rieffel, Optimizing quantum heuristics with meta-learning, Quantum Machine Intelli- gence 3, 13 (2021)
2021
-
[43]
Khairy, R
S. Khairy, R. Shaydulin, L. Cincio, Y. Alexeev, and P. Balaprakash, Learning to optimize variational quan- tum circuits to solve combinatorial problems, in Pro- ceedings of the AAAI conference on artificial intelligence, Vol. 34 (2020) pp. 2367–2375
2020
-
[44]
Li and J
K. Li and J. Malik, Learning to optimize, in International Conference on Learning Representations (2017)
2017
-
[45]
Y. Chen, M. W. Hoffman, S. G. Colmenarejo, M. Denil, T. P. Lillicrap, M. Botvinick, and N. Freitas, Learning to learn without gradient descent by gradient descent, in International Conference on Machine Learning (PMLR,
-
[46]
A. W. R. Smith, J. Gray, and M. S. Kim, Efficient quan- tum state sample tomography with basis-dependent neu- ral networks, PRX Quantum 2, 020348 (2021)
2021
-
[47]
Y. Quek, S. Fort, and H. K. Ng, Adaptive quantum state tomography with neural networks, npj Quantum Infor- 14 mation 7, 105 (2021)
2021
-
[48]
Kandala, A
A. Kandala, A. Mezzacapo, K. Temme, M. Takita, M. Brink, J. M. Chow, and J. M. Gambetta, Hardware- efficient variational quantum eigensolver for small molecules and quantum magnets, Nature 549, 242 (2017)
2017
-
[49]
Salimans, J
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever, Evolution strategies as a scalable alternative to reinforce- ment learning, arXiv preprint arXiv:1703.03864 (2017)
2017 arXiv
-
[50]
Narvekar, B
S. Narvekar, B. Peng, M. Leonetti, J. Sinapov, M. E. Taylor, and P. Stone, Curriculum learning for reinforce- ment learning domains: a framework and survey, J. Mach. Learn. Res. 21 (2020)
2020
-
[51]
Ostaszewski, L
M. Ostaszewski, L. M. Trenkwalder, W. Masarczyk, E. Scerri, and V. Dunjko, Reinforcement learning for op- timization of variational quantum circuit architectures, Advances in Neural Information Processing Systems 34, 18182 (2021)
2021
-
[52]
Y. J. Patel, A. Kundu, M. Ostaszewski, X. Bonet- Monroig, V. Dunjko, and O. Danaci, Curriculum rein- forcement learning for quantum architecture search un- der hardware errors, in The Twelfth International Con- ference on Learning Representations (2024)
2024
-
[53]
J. Haah, A. W. Harrow, Z. Ji, X. Wu, and N. Yu, Sample- optimal tomography of quantum states, IEEE Transac- tions on Information Theory 63, 5628 (2017)
2017
-
[54]
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction (MIT press, 2018)
2018
-
[55]
Bellman, The theory of dynamic programming, Bul- letin of the American Mathematical Society 60, 503 (1954)
R. Bellman, The theory of dynamic programming, Bul- letin of the American Mathematical Society 60, 503 (1954)
1954
-
[56]
K. J. ˚Astr¨ om, Optimal control of markov processes with incomplete state information ii: The convexity of the loss- function, Journal of Mathematical Analysis and Applica- tions 26, 403 (1969)
1969
-
[57]
Hausknecht and P
M. Hausknecht and P. Stone, Deep recurrent q-learning for partially observable mdps, in 2015 aaai fall sympo- sium series (2015)
2015
-
[58]
Barry, D
J. Barry, D. T. Barry, and S. Aaronson, Quantum par- tially observable markov decision processes, Phys. Rev. A 90, 032311 (2014)
2014
-
[59]
V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, and M. H. Devoret, Model-free quantum control with re- inforcement learning, Phys. Rev. X 12, 011059 (2022)
2022
-
[60]
Bukov, A
M. Bukov, A. G. R. Day, D. Sels, P. Weinberg, A. Polkovnikov, and P. Mehta, Reinforcement learning in different phases of quantum control, Phys. Rev. X 8, 031086 (2018)
2018
-
[61]
F¨ osel, P
T. F¨ osel, P. Tighineanu, T. Weiss, and F. Marquardt, Re- inforcement learning with neural networks for quantum feedback, Phys. Rev. X 8, 031084 (2018)
2018
-
[62]
M. Y. Niu, S. Boixo, V. N. Smelyanskiy, and H. Neven, Universal quantum control through deep reinforcement learning, npj Quantum Information 5, 33 (2019)
2019
-
[63]
Z. T. Wang, Y. Ashida, and M. Ueda, Deep reinforcement learning control of quantum cartpoles, Phys. Rev. Lett. 125, 100401 (2020)
2020
-
[64]
Porotti, A
R. Porotti, A. Essig, B. Huard, and F. Marquardt, Deep Reinforcement Learning for Quantum State Preparation with Weak Nonlinear Measurements, Quantum 6, 747 (2022)
2022
-
[65]
Metz and M
F. Metz and M. Bukov, Self-correcting quantum many- body control using reinforcement learning with tensor networks, Nature Machine Intelligence 5, 780 (2023)
2023
-
[66]
Vent, Rechenberg, ingo, evolutionsstrategie — opti- mierung technischer systeme nach prinzipien der biologis- chen evolution
W. Vent, Rechenberg, ingo, evolutionsstrategie — opti- mierung technischer systeme nach prinzipien der biologis- chen evolution. 170 s. mit 36 abb. frommann-holzboog- verlag. stuttgart 1973. broschiert, Feddes Repertorium 86, 337 (1975)
1975
-
[67]
Hansen, The cma evolution strategy: A tutorial, arXiv preprint arXiv:1604.00772 (2016)
N. Hansen, The cma evolution strategy: A tutorial, arXiv preprint arXiv:1604.00772 (2016)
2016 arXiv
-
[68]
Mitarai, M
K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, Quan- tum circuit learning, Phys. Rev. A 98, 032309 (2018)
2018
-
[69]
Schuld, V
M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Kil- loran, Evaluating analytic gradients on quantum hard- ware, Phys. Rev. A 99, 032331 (2019)
2019
-
[70]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Bab- bush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nature Communications 9, 4812 (2018)
2018
-
[71]
Cerezo and P
M. Cerezo and P. J. Coles, Higher order derivatives of quantum neural networks with barren plateaus, Quan- tum Science and Technology 6, 035006 (2021)
2021
-
[72]
Arrasmith, M
A. Arrasmith, M. Cerezo, P. Czarnik, L. Cincio, and P. J. Coles, Effect of barren plateaus on gradient-free opti- mization, Quantum 5, 558 (2021)
2021
-
[73]
Anand, M
A. Anand, M. Degroote, and A. Aspuru-Guzik, Natu- ral evolutionary strategies for variational quantum com- putation, Machine Learning: Science and Technology 2, 045012 (2021)
2021
-
[74]
J. Xie, C. Xu, C. Yin, Y. Dong, and Z. Zhang, Nat- ural evolutionary gradient descent strategy for varia- tional quantum algorithms, Intelligent Computing 2, 0042 (2023)
2023
-
[75]
Mezzadri, How to generate random matrices from the classical compact groups (2007), arXiv:math-ph/0609050 [math-ph]
F. Mezzadri, How to generate random matrices from the classical compact groups (2007), arXiv:math-ph/0609050 [math-ph]
2007 arXiv
-
[76]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al. , Pytorch: An imperative style, high-performance deep learning library, Advances in neural information processing systems 32 (2019)
2019
-
[77]
See https://github.com/quantum-jwjae/RL2LQS
-
[78]
S. J. Pan and Q. Yang, A survey on transfer learning, IEEE Transactions on Knowledge and Data Engineering 22, 1345 (2010)
2010
-
[79]
R. Zen, L. My, R. Tan, F. H´ ebert, M. Gattobigio, C. Miniatura, D. Poletti, and S. Bressan, Transfer learn- ing for scalability of neural-network quantum states, Phys. Rev. E 101, 053301 (2020)
2020
-
[80]
A. Mari, T. R. Bromley, J. Izaac, M. Schuld, and N. Killo- ran, Transfer learning in hybrid classical-quantum neural networks, Quantum 4, 340 (2020)
2020
-
[81]
Bagan, M
E. Bagan, M. Baig, R. Mu˜ noz Tapia, and A. Rodriguez, Collective versus local measurements in a qubit mixed- state estimation, Phys. Rev. A 69, 010304 (2004)
2004
-
[82]
D. H. Mahler, L. A. Rozema, A. Darabi, C. Ferrie, R. Blume-Kohout, and A. M. Steinberg, Adaptive quan- tum state tomography improves accuracy quadratically, Phys. Rev. Lett. 111, 183601 (2013)
2013
-
[83]
K. S. Kravtsov, S. S. Straupe, I. V. Radchenko, N. M. T. Houlsby, F. Husz´ ar, and S. P. Kulik, Experimental adap- tive bayesian tomography, Phys. Rev. A 87, 062122 (2013)
2013
-
[84]
Ferrie, Self-guided quantum tomography, Phys
C. Ferrie, Self-guided quantum tomography, Phys. Rev. Lett. 113, 190404 (2014)
2014
-
[85]
Shen and S
J. Shen and S. Castan, An optimal linear operator for step edge detection, CVGIP: Graphical Models and Im- 15 age Processing 54, 112 (1992)
1992
-
[86]
D. N. Page, Average entropy of a subsystem, Phys. Rev. Lett. 71, 1291 (1993)
1993
-
[87]
Sharma, A
S. Sharma, A. S. Lakshminarayanan, and B. Ravindran, Learning to repeat: Fine grained action repetition for deep reinforcement learning, in International Conference on Learning Representations (2017)
2017
-
[88]
Silver, A
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, Mas...
2016
-
[89]
Porotti, V
R. Porotti, V. Peano, and F. Marquardt, Gradient- ascent pulse engineering with feedback, PRX Quantum 4, 030305 (2023)
2023
-
[90]
I. S. Maria Schuld and F. Petruccione, An introduction to quantum machine learn- ing, Contemporary Physics 56, 172 (2015), https://doi.org/10.1080/00107514.2014.964942
2015
-
[91]
Abbas, A
A. Abbas, A. Ambainis, B. Augustino, A. B¨ artschi, H. Buhrman, C. Coffrin, G. Cortiana, V. Dunjko, D. J. Egger, B. G. Elmegreen, N. Franco, F. Fratini, B. Fuller, J. Gacon, C. Gonciulea, S. Gribling, S. Gupta, S. Had- field, R. Heese, G. Kircher, T. Kleinert, T. Koch, G. Ko- ...
2024
-
[92]
Y. Chen, M. W. Hoffman, S. G. Colmenarejo, M. Denil, T. P. Lillicrap, M. Botvinick, and N. de Freitas, Learn- ing to learn without gradient descent by gradient de- scent, in Proceedings of the 34th International Conference on Machine Learning , Proceedings of Machine Learning ...
2017
-
[93]
V. TV, P. Malhotra, J. Narwariya, L. Vig, and G. Shroff, Meta-learning for black-box optimization, in Joint Eu- ropean Conference on Machine Learning and Knowledge Discovery in Databases (Springer, 2019) pp. 366–381
2019
-
[94]
Shala, A
G. Shala, A. Biedenkapp, N. Awad, S. Adriaensen, M. Lindauer, and F. Hutter, Learning step-size adapta- tion in cma-es, in Parallel Problem Solving from Nature – PPSN XVI , edited by T. B¨ ack, M. Preuss, A. Deutz, H. Wang, C. Doerr, M. Emmerich, and H. Trautmann (Springer Int...
2020
-
[95]
R. T. Lange, T. Schaul, Y. Chen, T. Zahavy, V. Dal- ibard, C. Lu, S. Singh, and S. Flennerhag, Discovering evolution strategies via meta-black-box optimization, in The Eleventh International Conference on Learning Rep- resentations (2023)
2023
-
[96]
Konda and J
V. Konda and J. Tsitsiklis, Actor-critic algorithms, Ad- vances in Neural Information Processing Systems 12 (1999)
1999
-
[97]
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. De Freitas, Sample efficient actor-critic with experience replay, Proceedings of the International Conference on Learning Representations (ICLR) (2017)
2017
-
[98]
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beat- tie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, Human-level con- trol through de...
2015
-
[99]
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, Asynchronous methods for deep reinforcement learning, in Proceedings of The 33rd International Conference on Machine Learn- ing, Proceedings of Machine Learning Research, Vol. 48...
2016
-
[100]
Kocsis and C
L. Kocsis and C. Szepesv´ ari, Bandit based monte-carlo planning, in Machine Learning: ECML 2006 , edited by J. F¨ urnkranz, T. Scheffer, and M. Spiliopoulou (Springer Berlin Heidelberg, Berlin, Heidelberg, 2006) pp. 282–293
2006
-
[101]
D. P. Kingma and J. Ba, Adam: A method for stochastic optimization, arXiv preprint arXiv:1412.6980 (2014)
2014 arXiv
-
[102]
ˇReh´ aˇ cek, Z
J. ˇReh´ aˇ cek, Z. c. v. Hradil, E. Knill, and A. I. Lvovsky, Diluted maximum-likelihood algorithm for quantum to- mography, Phys. Rev. A 75, 042108 (2007)
2007
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.