REVIEW 4 major objections 5 minor 71 references
Quantum Many-body Simulations from a Reinforcement-Learned Exponential Ansatz
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper claims that a reinforcement-learned agent can select the few two-body exponential transformations that reach chemical accuracy for H3 and H4, using five to ten actions where a filtered standard solver needs over twenty.
desk verdict A credible, useful incremental step that compresses CQE circuits with RL on small molecules, but the paper overstates the Markovian state claim and needs statistical and baseline strengthening before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the Markovian decision process whose state is the CSE residual operator $\hat{R} = \sum_{ijkl} {}^2R^{ij}_{kl}\hat{\Gamma}^{ij}_{kl}$ and whose actions are the individual two-body exponential factors $e^{\theta \hat{\Gamma}_n}$ from the Trotterized product ansatz. The agent is a dueling double deep Q-network: two streams estimate state value and per-action advantage, trained with prioritized experience replay, a reuse penalty to discourage repeating operators, and a line-search over the continuous $\theta$ embedded in the discrete action. The paper's key comparison is against filtered CQE, which keeps only the five largest residual coefficients per iteration; the RL agent outperforms it because it can merge or reorder effectively commuting operators that filter-based selection applies redundantly.
What would settle it
Train the same RL-CQE protocol on a molecule whose ground state is strongly correlated beyond two-body information—for example, a stretched chain of four or six hydrogen atoms—and check whether a five-to-ten-action policy still reaches chemical accuracy. If the optimal next operator depends on information not captured by the two-electron residual, the validation error will rise above the ~1.6 mHa chemical threshold, directly testing the Markovian state assumption.
Extended reading notes
Core claim
The central claim is that the contracted Schrödinger equation (CSE) residual—the two-electron projection of the Schrödinger equation at the current state—is a complete state description for choosing the next wavefunction update, making ansatz construction a Markovian decision process. In this picture the agent's action is a single two-body exponential factor $e^{\theta \hat{\Gamma}_n}$ from the exact universal two-body exponential ansatz, and its reward is the negative of energy plus a weighted residual norm, $Q = -(E + \lambda \|R\|)$. A dueling double deep Q-network learns to select actions from the current residual; because the residual already encodes the Hamiltonian through commutation, the learned policy transfers across molecular geometries. The demonstrated consequence is that near-optimal ansätze of five to ten exponentials achieve chemical accuracy for H3 and H4, outperforming the filtered CQE baseline and suggesting a practical strategy for resource-limited quantum simulation.
Load-bearing premise
The load-bearing premise is that the two-electron CSE residual at the current wavefunction contains enough information to pick the optimal next exponential transformation; if deciding that step ever requires higher-order reduced density matrices or a memory of past actions, the Markovian state is incomplete and the learned compaction could fail on larger molecules.
Editorial extensions
If this is right
- Chemical accuracy for H3 and H4 with five to ten two-body exponential transformations, so RL-CQE circuits are roughly two to four times shallower than filtered CQE.
- Policies trained on a distribution of molecular geometries transfer to unseen geometries, with validation energy errors of 1.2–2.4 mHa, suggesting geometry-robust ansatz generation.
- Because the state is the CSE residual, the learned policy may transfer across Hamiltonians and devices, including training on noisy quantum hardware where depth constraints are strict.
- The RL formulation extends beyond CQE: the same Markovian action framework can be applied to other parameterized quantum eigensolvers such as VQE to optimize circuit architecture.
Reading between the lines
- If the Markovian assumption survives scaling, RL-CQE could become an automated ansatz designer: training once on a distribution of Hamiltonians and then deploying the policy on a quantum device would eliminate per-molecule circuit design and classical optimization loops.
- The reuse penalty hints that the agent learns to avoid repeated operators, effectively discovering commutativity relations; a natural test is whether policies trained on H3 transfer to H4 or to larger basis sets, which the paper does not report.
- Pairing the CSE residual state with a transformer-based policy (as the authors mention in outlook) would let the agent see operator histories, potentially correcting any non-Markovian effects and extending the method to molecules with stronger multi-reference character.
- The depth–accuracy trade-off exposed by varying allowed actions (3, 5, 10) could be mapped against noise models, letting users pick the minimal circuit that achieves a target accuracy on a specific device — a resource-allocation tool for near-term quantum simulation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces a reinforcement learning (RL) treatment of the contracted quantum eigensolver (CQE). The central idea is to formulate the iterative CQE wavefunction update as a Markov decision process in which the state is the two-electron contracted Schrödinger equation (CSE) residual, the actions are two-body exponential transformations of the wavefunction, and a dueling double deep Q-network is trained to select actions that minimize both energy and residual norm. The authors report that for H3 and H4 in the STO-3G basis, the RL-optimized ansatz reaches chemical accuracy with 5–10 exponential transformations, whereas filtered CQE requires more than 20, and that a policy trained on a grid of H3 geometries transfers to held-out geometries with energy errors of 1.2–2.4 mHa. They also discuss transfer between classical simulation and quantum hardware and provide open-source code.
Significance. If the claims hold, the paper makes a useful contribution: it connects CQE's exact two-body exponential ansatz with modern deep RL, provides an open-source implementation, and proposes a concrete accuracy-versus-circuit-depth benchmark. The empirical results are suggestive but preliminary: they involve two small molecules, single runs without statistical error bars, and validation errors that exceed the standard chemical-accuracy threshold for most folds. The main theoretical premise—that the CSE residual is a Markovian and sufficient state for action selection—is asserted rather than proved and is questionable in light of Eq. (6). The broader transferability claims therefore need additional empirical or theoretical support.
major comments (4)
- [Sec. II.B, item (1) and Eq. (6)] The paper asserts that the CSE residual provides a two-electron description of the current state and that the update from S_t(^2D) to S_{t+1}(^2D) is Markovian. However, Eq. (6) defines 2R^{ij}_{kl} = <Ψ|Γ^{ij}_{kl}(H−E)|Ψ>. Because H contains one- and two-body terms and Γ contains four fermionic operators, this expectation value depends on the 3- and 4-RDMs, not only on the 2-RDM. Two wavefunctions with identical 2-RDMs can therefore have different residuals and, consequently, different optimal next exponential actions. The current residual does not determine the next residual after an action, so the MDP state is not Markovian as stated. This is load-bearing for the claim that the learned policy transfers to quantum-hardware settings where only residual measurements are available. Please either prove the Markovian property under the exact classical update, or reframe the state as a heuristic sufficient statistic and provide empirical tests (for example, compare policies trained with residual-only inputs against policies trained with residual plus action history or additional RDM information). The authors' own Sec. IV remark that moving beyond the Markovian process is future work is consistent with this gap.
- [Sec. III.B, Fig. 4] The key quantitative comparison between RL-optimized and filtered CQE ansätze is presented as single convergence curves without error bars or multiple random seeds. As a result, the claim that RL reaches chemical accuracy with 5–10 actions while filtered CQE requires more than 20 is not statistically substantiated. Please provide means and standard deviations over at least 5–10 independent training runs, and explicitly state the chemical-accuracy threshold (presumably 1.6 mHa) used to define 'chemical accuracy' in Fig. 4.
- [Sec. III.B, second paragraph] The text notes that for H4 with five actions the agent can only converge the residual to less than 0.007, suggesting that the solution may not be at chemical accuracy. This is important for the interpretation of Fig. 2 and the depth-versus-accuracy trade-off. Please report the corresponding energy errors for the five-action H4 runs and clarify how the residual norm relates to the energy accuracy criterion.
- [Table I and Sec. III.C] The validation energy differences in Table I range from 1.21 to 2.44 mHa, i.e., above the standard 1.6 mHa chemical-accuracy threshold for most folds. The statement that the agent 'performs well in predicting the validation set' therefore needs qualification. Please report the fraction of held-out geometries within chemical accuracy, include standard deviations across the five folds, and provide a comparison with unfiltered CQE and with a standard ADAPT-VQE construction using the same operator pool to support the claimed advantage over conventional ansätze mentioned in Sec. III.C.
minor comments (5)
- The reward function in Eq. (8) does not include the penalty for reused actions described in the text; please define the full reward including the reuse penalty.
- The product over non-commuting exponentials in Eq. (7) is not specified; please state the ordering convention (e.g., right-to-left time ordering or a fixed sequence).
- Please provide the full set of DQN hyperparameters (number of episodes, replay buffer size, batch size, target-network update period, epsilon schedule, and any reward shaping constants) in a table or appendix for reproducibility; the text currently lists only some of them.
- The claim that using the residual as the state 'greatly expands the model's transferability' is not tested against an alternative state representation (e.g., the raw 2-RDM or energy plus 2-RDM); an ablation would strengthen this claim.
- The phrase 'we reformulate the wavefunction update as a Markovian decision process' is stronger than what is demonstrated; consider changing to 'treat as' or providing a proof of the Markov property.
Circularity Check
No significant circularity: the RL-CQE results are empirical benchmarks and the ansatz exactness is imported from independent published proofs, not from the present fitting procedure.
full rationale
The central claims—that RL can select compact sequences of two-body exponentials and that the trained policy transfers to unseen geometries—are empirical results benchmarked against exact diagonalization, not outputs of the ansatz exactness theorems. The exactness of the two-body exponential ansatz and the CSE equivalence are imported from published proofs (Refs. [23,37,38,60,61]), including independent proof by Evangelista et al. (Ref. [61]); these are parameter-free mathematical results with stated assumptions that do not include the RL target, so they are independent support rather than self-citation circularity. The reward is defined directly from energy and residual norm and is minimized against exact answers; no fitted parameter is renamed as a prediction. Transfer experiments use held-out geometries, so validation errors are genuine out-of-sample predictions. The one asserted-but-unproven premise is the Markovian sufficiency of the CSE residual as an RL state (Sec. II.B). Equation (6) shows the residual tensor depends on higher-body reduced density matrices, so the claim that the update from the 2-RDM alone is Markovian is not established; however, this is an unsupported assumption or correctness risk, not a reduction of a prediction to its input, and the paper itself flags moving beyond the Markovian process as future work in Sec. IV. No circular step is exhibited.
Assumptions & free parameters
free parameters (6)
- reward weight lambda =
0.2
- maximum number of actions per episode =
3, 5, 10, 20 (tested)
- DQN discount factor =
0.99
- replay buffer priority exponent alpha =
0.6
- epsilon-greedy decay rate =
0.999 per episode
- target network update period =
1000 episodes
assumptions (5)
- domain assumption A wavefunction satisfies the CSE if and only if it satisfies the Schrödinger equation.
- domain assumption Any N-electron wavefunction can be represented exactly by a product of two-body exponential transformations.
- domain assumption The CSE residual at the current iteration is a sufficient, Markovian state representation for choosing the next exponential transformation.
- ad hoc to paper The reward Q = -(E + lambda*||R||) is an adequate proxy for accuracy and circuit-depth trade-off.
- domain assumption Trotterization error of the product of exponentials is negligible for the convergence claims.
Cite this review
Pith. "Pith review of Quantum Many-body Simulations from a Reinforcement-Learned Exponential Ansatz." pith.science (2026). https://pith.science/paper/HEOZHE7K
@misc{pith2026250501935,
author = {Pith},
title = {Pith review of: Quantum Many-body Simulations from a Reinforcement-Learned Exponential Ansatz},
year = {2026},
howpublished = {\url{https://pith.science/paper/HEOZHE7K}},
note = {Machine review of arXiv:2505.01935}
}
read the original abstract
Solving for the many-body wavefunction represents a significant challenge on both classical and quantum devices because of the exponential scaling of the Hilbert space with system size. While the complexity of the wavefunction can be reduced through conventional ans\"{a}tze (e.g., the coupled cluster ansatz), it can still grow rapidly with system size even on quantum devices. An exact, universal two-body exponential ansatz for the many-body wavefunction has been shown to be generated from the solution of the contracted Schr\"odinger equation (CSE), and recently, this ansatz has been implemented without classical approximation on quantum simulators and devices for the scalable simulation of many-body quantum systems. Here we combine the solution of the CSE with a form of artificial intelligence known as reinforcement learning (RL) to generate highly compact circuits that implement this ansatz without sacrificing accuracy. As a natural extension of CSE, we reformulate the wavefunction update as a Markovian decision process and train the agent to select the optimal actions at each iteration based upon only the current CSE residual. Compact circuits with high accuracy are achieved for H3 and H4 molecules over a range of molecular geometries.
Figures
Reference graph
Works this paper leans on
-
[1]
B. P. Lanyon, J. D. Whitfield, G. G. Gillett, M. E. Goggin, M. P. Almeida, I. Kassal, J. D. Biamonte, M. Mohseni, B. J. Powell, M. Barbieri, et al. , Towards quantum chemistry on a quantum computer, Nat. Chem. 2, 106 (2010)
work page 2010
-
[2]
P. J. O’Malley, R. Babbush, I. D. Kivlichan, J. Romero, J. R. McClean, R. Barends, J. Kelly, P. Roushan, A. Tranter, N. Ding, et al. , Scalable quantum simulation of molecular energies, Phys. Rev. X 6, 031007 (2016)
work page 2016
-
[3]
McArdle, S
S. McArdle, S. Endo, A. Aspuru-Guzik, S. C. Benjamin, and X. Yuan, Quantum computational chemistry, Rev. Mod. Phys. 92, 015003 (2020)
2020
- [4]
-
[5]
Cerezo, A
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, et al. , Variational quantum algorithms, Nat. Rev. Phys. 3, 625 (2021)
2021
-
[6]
Kandala, A
A. Kandala, A. Mezzacapo, K. Temme, M. Takita, M. Brink, J. M. Chow, and J. M. Gambetta, Hardware- efficient variational quantum eigensolver for small molecules and quantum magnets, nature549, 242 (2017)
2017
- [7]
-
[8]
Wecker, M
D. Wecker, M. B. Hastings, and M. Troyer, Progress to- wards practical quantum variational algorithms, Phys. Rev. A 92, 042303 (2015)
2015
Show all 71 references
-
[9]
D’Cunha, T
R. D’Cunha, T. D. Crawford, M. Motta, and J. E. Rice, Challenges in the use of quantum computing hardware- efficient ansa¨tze in electronic structure theory, J. Phys. Chem. A 127, 3437 (2023)
2023
- [10]
-
[11]
S. E. Smart and D. A. Mazziotti, Quantum solver of con- tracted eigenvalue equations for scalable molecular simu- lations on quantum computing devices, Phys. Rev. Lett. 126, 070504 (2021)
2021
-
[12]
S. E. Smart and D. A. Mazziotti, Accelerated con- vergence of contracted quantum eigensolvers through a quasi-second-order, locally parameterized optimization, J. Chem. Theory Comput. 18, 5286 (2022)
2022
-
[13]
J.-N. Boyn, A. O. Lykhin, S. E. Smart, L. Gagliardi, and D. A. Mazziotti, Quantum-classical hybrid algorithm for the simulation of all-electron correlation, J. Chem. Phys. 155, 244106 (2021)
2021
-
[14]
S. E. Smart, J.-N. Boyn, and D. A. Mazziotti, Resolving correlated states of benzyne with an error-mitigated con- tracted quantum eigensolver, Phys. Rev. A 105, 022405 (2022)
2022
-
[15]
Y. Wang, L. M. Smith, and D. A. Mazziotti, Quantum simulation of bosons with the contracted quantum eigen- solver, New J. Phys. 25, 103005 (2023)
2023
-
[16]
Wang and D
Y. Wang and D. A. Mazziotti, Electronic excited states from a variance-based contracted quantum eigensolver, Phys. Rev. A 108, 022814 (2023)
2023
-
[17]
C. L. Benavides-Riveros, Y. Wang, S. Warren, and D. A. Mazziotti, Quantum simulation of excited states from parallel contracted quantum eigensolvers, New J. Phys. 26, 033020 (2024)
2024
-
[18]
S. E. Smart and D. A. Mazziotti, Verifiably exact solu- tion of the electronic Schr¨ odinger equation on quantum devices, Phys. Rev. A 109, 022802 (2024)
2024
-
[19]
S. E. Smart, D. M. Welakuh, and P. Narang, Many-Body Excited States with a Contracted Quantum Eigensolver, J. Chem. Theory Comput. 20, 3580 (2024), 2305.09653
2024 arXiv
-
[20]
Warren, Y
S. Warren, Y. Wang, C. L. Benavides-Riveros, and D. A. Mazziotti, Exact ansatz of fermion-boson systems for a quantum device, Phys. Rev. Lett. 133, 080202 (2024)
2024
-
[21]
Warren, Y
S. Warren, Y. Wang, C. L. Benavides-Riveros, and D. A. Mazziotti, Quantum algorithm for polaritonic chemistry based on an exact ansatz, Quantum Sci. Technol. 10, 02LT02 (2025)
2025
-
[22]
Y. Wang, C. Cianci, I. Avdic, R. Dutta, S. Warren, B. Allen, N. P. Vu, L. F. Santos, V. S. Batista, and D. A. Mazziotti, Characterizing conical intersections of nucle- obases on quantum computers, J. Chem. Theory Com- put. 10.1021/acs.jctc.4c01434
-
[23]
D. A. Mazziotti, Contracted Schr¨ odinger equation: De- termining quantum energies and two-particle density ma- trices without wave functions, Phys. Rev. A 57, 4219 (1998)
1998
-
[24]
Colmenero and C
F. Colmenero and C. Valdemoro, Approximating q-order reduced density matrices in terms of the lower-order ones. II. Applications, Phys. Rev. A 47, 979 (1993)
1993
-
[25]
Nakatsuji and K
H. Nakatsuji and K. Yasuda, Direct Determination of the Quantum-Mechanical Density Matrix Using the Density Equation, Phys. Rev. Lett. 76, 1039 (1996)
1996
-
[26]
D. A. Mazziotti, Comparison of contracted Schr¨ odinger and coupled-cluster theories, Phys. Rev. A 60, 4396 (1999). 7
1999
-
[27]
Mukherjee and W
D. Mukherjee and W. Kutzelnigg, Irreducible Brillouin conditions and contracted Schr¨ odinger equations for n- electron systems. I. The equations satisfied by the density cumulants, J. Chem. Phys. 114, 2047 (2001)
2001
-
[28]
Yasuda, Uniqueness of the solution of the contracted Schr¨ odinger equation, Phys
K. Yasuda, Uniqueness of the solution of the contracted Schr¨ odinger equation, Phys. Rev. A65, 052121 (2002)
2002
-
[29]
D. A. Mazziotti, Variational method for solving the con- tracted Schr¨ odinger equation through a projection of the N-particle power method onto the two-particle space, J. Chem. Phys. 116, 1239 (2002)
2002
-
[30]
Cohen and C
L. Cohen and C. Frishberg, Hierarchy Equations for Re- duced Density Matrices, Phys. Rev. A 13, 927 (1976)
1976
-
[31]
Valdemoro, L
C. Valdemoro, L. Tel, D. Alcoba, and E. P´ erez-Rome- ro, The contracted Schr¨ odinger equation methodology: study of the third-order correlation effects, Theor. Chem. Account. 118, 503 (2007)
2007
-
[32]
D. A. Mazziotti, Anti-Hermitian Contracted Schr¨ odin- ger Equation: Direct Determination of the Two-Electron Reduced Density Matrices of Many-Electron Molecules, Phys. Rev. Lett. 97, 143002 (2006)
2006
-
[33]
D. A. Mazziotti, Anti-Hermitian part of the contracted Schr¨ odinger equation for the direct calculation of two- electron reduced density matrices, Phys. Rev. A 75, 022505 (2007)
2007
-
[34]
D. A. Mazziotti, Multireference many-electron correla- tion energies from two-electron reduced density matri- ces computed by solving the anti-Hermitian contracted Schr¨ odinger equation, Phys. Rev. A76, 052502 (2007)
2007
-
[35]
Boyn and D
J.-N. Boyn and D. A. Mazziotti, Accurate singlet–triplet gaps in biradicals via the spin averaged anti-Hermitian contracted Schr¨ odinger equation, J. Chem. Phys. 154, 134103 (2021)
2021
-
[36]
Bonet-Monroig, R
X. Bonet-Monroig, R. Babbush, and T. E. O’Brien, Nearly optimal measurement scheduling for partial to- mography of quantum states, Phys. Rev. X 10, 031064 (2020)
2020
-
[37]
D. A. Mazziotti, Exactness of wave functions from two- body exponential transformations in many-body quan- tum theory, Phys. Rev. A 69, 012507 (2004)
2004
-
[38]
D. A. Mazziotti, Exact two-body expansion of the many- particle wave function, Phys. Rev. A 102, 030802 (2020)
2020
-
[39]
Gidofalvi and D
G. Gidofalvi and D. A. Mazziotti, Direct calculation of excited-state electronic energies and two-electron re- duced density matrices from the anti-Hermitian con- tracted Schr¨ odinger equation, Phys. Rev. A 80, 022507 (2009)
2009
-
[40]
A. E. Rothman, J. J. Foley, and D. A. Mazziotti, Open- shell energies and two-electron reduced density matrices from the anti-Hermitian contracted Schr¨ odinger equa- tion: A spin-coupled approach, Phys. Rev. A 80, 052508 (2009)
2009
-
[41]
J. W. Snyder and D. A. Mazziotti, Photoexcited con- version of gauche-1,3-butadiene to bicyclobutane via a conical intersection: Energies and reduced density matri- ces from the anti-Hermitian contracted Schr¨ odinger equa- tion, J. Chem. Phys. 135, 024107 (2011)
2011
-
[42]
A. M. Sand and D. A. Mazziotti, Enhanced computa- tional efficiency in the direct determination of the two- electron reduced density matrix from the anti-Hermitian contracted Schr¨ odinger equation with application to ground and excited states of conjugated π-systems, J. Chem....
2015
-
[43]
L. P. Kaelbling, M. L. Littman, and A. W. Moore, Re- inforcement Learning: A Survey, J. Artif. Intell. Res. 4, 237 (1996)
1996
-
[44]
H. V. Hasselt, A. Guez, and D. Silver, Deep Reinforce- ment Learning with Double Q-Learning, Proceedings of the AAAI Conference on Artificial Intelligence 30, 10.1609/aaai.v30i1.10295 (2016)
2016 doi
-
[45]
Silver, T
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play., Science (New York,...
2018
-
[46]
Vinyals, I
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard,...
2019
-
[47]
M. Y. Niu, S. Boixo, V. N. Smelyanskiy, and H. Neven, Universal quantum control through deep reinforcement learning, npj Quantum Inf. 5, 33 (2019)
2019
-
[48]
Liang, L
S. Liang, L. Zhu, X. Liu, C. Yang, and X. Li, Artificial- intelligence-driven shot reduction in quantum measure- ment, Chem. Phys. Rev. 5, 10.1063/5.0219663 (2024)
2024 doi
-
[49]
Ostaszewski, L
M. Ostaszewski, L. M. Trenkwalder, W. Masarczyk, E. Scerri, and V. Dunjko, Reinforcement learning for optimization of variational quantum circuit architec- tures, in Advances in Neural Information Processing Sys- tems, Vol. 34, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin,...
2021
-
[50]
Zhang, P.-L
Y.-H. Zhang, P.-L. Zheng, Y. Zhang, and D.-L. Deng, Topological quantum compiling with reinforce- ment learning, Phys. Rev. Lett. 125, 170501 (2020)
2020
-
[51]
H. P. Nautrup, N. Delfosse, V. Dunjko, H. J. Briegel, and N. Friis, Optimizing quantum error correction codes with reinforcement learning, Quantum 3, 215 (2019)
2019
-
[52]
Andreasson, J
P. Andreasson, J. Johansson, S. Liljestrand, and M. Granath, Quantum error correction for the toric code using deep reinforcement learning, Quantum 3, 183 (2019)
2019
-
[53]
J. Olle, R. Zen, M. Puviani, and F. Marquardt, Si- multaneous discovery of quantum error correction codes and encoders with a noise-aware reinforcement learning agent, Npj Quantum Inf. 10, 1 (2024)
2024
-
[54]
Fallani, M
A. Fallani, M. A. Rossi, D. Tamascelli, and M. G. Genoni, Learning feedback control strategies for quantum metrol- ogy, PRX Quantum 3, 020310 (2022)
2022
-
[55]
Belliardo, F
F. Belliardo, F. Zoratti, F. Marquardt, and V. Gio- vannetti, Model-aware reinforcement learning for high- performance bayesian experimental design in quantum metrology, Quantum 8, 1555 (2024)
2024
-
[56]
A. J. Coleman, Structure of Fermion Density Matrices, Rev. Mod. Phys. 35, 668 (1963)
1963
-
[57]
Yokonuma, Tensor Spaces and Exterior Alge- bra, Translations of Mathematical Monographs 10.1090/mmono/108 (1992)
T. Yokonuma, Tensor Spaces and Exterior Alge- bra, Translations of Mathematical Monographs 10.1090/mmono/108 (1992)
1992 doi
-
[58]
D. A. Mazziotti, Quantum Chemistry without Wave Functions: Two-Electron Reduced Density Matrices, Acc. Chem. Res. 39, 207 (2006). 8
2006
-
[59]
D. A. Mazziotti, Contracted Schr¨ odinger Equation, in Reduced-Density-Matrix Mechanics: With Application to Many-Electron Atoms and Molecules (John Wiley & Sons, Ltd, 2007) Chap. 8, pp. 165–203
2007
-
[60]
Nakatsuji, Equation for the Direct Determination of the Density Matrix, Phys
H. Nakatsuji, Equation for the Direct Determination of the Density Matrix, Phys. Rev. A 14, 41 (1976)
1976
-
[61]
F. A. Evangelista, G. K. Chan, and G. E. Scuseria, Exact parameterization of fermionic wave functions via unitary coupled cluster theory, J. Chem. Phys. 151, 10.1063/1.5133059 (2019)
2019 doi
-
[62]
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, Dueling network architectures for deep re- inforcement learning, in International conference on ma- chine learning (PMLR, 2016) pp. 1995–2003
2016
-
[63]
W. J. Hehre, R. F. Stewart, and J. A. Pople, Self- consistent molecular-orbital methods. i. use of gaussian expansions of slater-type atomic orbitals, J. Chem. Phys. 51, 2657 (1969)
1969
-
[64]
Jordan and E
P. Jordan and E. Wigner, ¨Uber das Paulische ¨Aquivalenzverbot, Z. Physik 47, 631 (1928)
1928
-
[65]
RDMChem, Quantum Chemistry Toolbox for Maple (Maplesoft, Waterloo, 2025)
2025
-
[66]
Maplesoft, Maple (Maplesoft, Waterloo, 2025)
2025
-
[67]
Wang and D
Y. Wang and D. A. Mazziotti, RL-CQE: A reinforce- ment learning approach to the contracted quantum eigen- solver, https://github.com/damazz/ML-CQE (2025), ac- cessed: May 3, 2025
2025
- [68]
-
[69]
Nakaji, L
K. Nakaji, L. B. Kristensen, J. A. Campos-Gonzalez- Angulo, M. G. Vakili, H. Huang, M. Bagherimehrab, C. Gorgulla, F. Wong, A. McCaskey, J.-S. Kim, et al. , The generative quantum eigensolver (gqe) and its application for ground state search, arXiv preprint arXiv:2401.09253 10...
2024 doi
- [70]
- [71]
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.