REVIEW 3 major objections 5 minor 123 references
Learning to stabilize nonequilibrium phases of matter with active feedback using partial information
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Reinforcement-learning agents trained on partial observations learn stochastic feedback policies that drive stabilizer circuits into area-law entangled steady states for any nonzero disentangling bias p, and above a critical information thr
desk verdict RL-trained feedback with partial information in Clifford circuits is a genuinely new mechanism; the finite-size evidence is solid, the thermodynamic-limit claim is a reasonable but unproven extrapolation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two ingredients carry the argument. First, a scalable active-feedback loop: at each turn the agent observes a (possibly partial) stabilizer tableau, outputs a bond location, and an optimally disentangling Clifford gate is applied there; the reward is negative normalized total entanglement, and the policy is a transformer encoder trained with Proximal Policy Optimization. Second, the mechanism the policy discovers: entanglement bottlenecks. By holding certain roughly equidistant bonds at near-zero entanglement, the agent splits the chain into weakly entangled subsystems, which appears in the steady state as self-similar pyramids in the bond-entanglement profile. A simplified rate model with n
What would settle it
Run the same learned-feedback training at N=256 for p=0.05 and inspect the steady state. The central claim predicts the pyramid count doubles to 8 and $\langle S_{\mathrm{tot}}\rangle/\mathcal{N}$ keeps decreasing roughly as $(1-p)/(pN)$. If the pyramid count saturates, the entanglement density plateaus, or a single-cut entropy at mid-chain grows linearly with N, the area-law claim fails.
Extended reading notes
Core claim
In the stabilizer-circuit model considered here, a bias p selects whether an optimally disentangling two-qubit Clifford gate or a random Clifford gate acts, while the RL agent chooses the location of the disentangling gate from a tableau observation. The central discovery is that learned policies solve the entanglement-suppression task in a way neither random nor greedy strategies do. With full information, the learned strategies eliminate the Clifford phase transition: for every fixed p>0, the normalized total entanglement $\langle S_{\mathrm{tot}}\rangle/\mathcal{N}$ decreases with system size and appears to vanish in the thermodynamic limit, so the volume-law phase survives only at p=0. W
Load-bearing premise
The area-law conclusion rests on the assumption that the self-similar pyramid pattern seen from N=64 to N=128 (pyramid count doubling from 2 to 4 at p=0.05) keeps doubling as N grows, so total entanglement grows only linearly and the normalized density continues to fall roughly as 1/(pN); only four system sizes, with visible deviations at p=0.05 and p=0.10, currently support that extrapolation.
Editorial extensions
If this is right
- If the area-law claim holds, the volume-law phase of this circuit model collapses to the single point p=0 for learned strategies: any nonzero disentangling bias keeps the steady state area-law.
- Partial information suffices only above q_c(p); the same RL pipeline reverts to volume-law entanglement below the threshold, and the required information decreases as p grows.
- Active feedback is essential: the learned strategy adapts in real time, whereas fixed human-designed bottleneck placement becomes unstable and reaches maximally entangled states at low p.
- The asymptotic law $\langle S_{\mathrm{tot}}\rangle/\mathcal{N}\sim(1-p)/(pN)$ gives a quantitative target for larger simulations and a discriminator between area-law scaling and finite-size volume-law behavior.
- The steady states are genuine nonequilibrium states: information-based control breaks detailed balance, and maintaining them carries a thermodynamic cost for erasing the controller's record.
Reading between the lines
- If the bottleneck mechanism is generic rather than Clifford-specific, the same learned feedback should suppress volume-law entanglement in Haar-random circuits, where the paper notes passive dynamics have no transition; moderate system sizes could already show pyramidal suppression.
- The conjecture that q=0 works for p at or above the greedy threshold implies that an experiment needs only to detect and target nearest-neighbor Bell-like pairs from single-shot measurements, not reconstruct the full quantum state, removing a major postselection bottleneck.
- The information threshold q_c(p) resembles a phase transition in the space of control policies; near it, RL training time or sample complexity may diverge with system size, making the threshold observable as a learning transition rather than only a steady-state property.
- A concrete prediction beyond the paper's data: at fixed small p, doubling N should double the number of pyramids and roughly halve the normalized entanglement, so N=256 at p=0.05 should show eight pyramids if the self-similar scaling continues.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies active feedback control of entanglement in (1+1)-dimensional stabilizer circuits using reinforcement learning. At each time step a disentangling two-qubit gate is placed according to a learned strategy π(a|o) based on partial observations of the stabilizer tableau, while random Clifford gates are applied with complementary probability. The authors compare random, greedy, and learned strategies and report that, for full information, learned strategies above a critical information threshold produce area-law steady-state entanglement for any bias p>0, eliminating the volume-law phase of the passive unitary-circuit game. They identify pyramid-shaped entanglement bottlenecks as the control mechanism and propose a simplified scaling model. For partial information, they report a critical information fraction qc(p) separating area-law from volume-law steady states. Simulations cover N=16,32,64,128, and code, data, and trained models are deposited.
Significance. If the thermodynamic-limit claim were established, the paper would demonstrate a new capability: information-driven reinforcement learning can engineer nonequilibrium steady states with qualitatively different entanglement scaling than passive dynamics, using only partial state information. The work is clearly presented, compares learned strategies against meaningful baselines, and identifies an interpretable mechanism (bottlenecks/pyramids) rather than a black-box policy. Strengths include the reproducible deposition of code, data, and trained models, the direct simulation of a well-defined stochastic process with no fitted parameters in the central results, and the detailed RL methodology. The main limitation is that the headline claim of area-law scaling for any p>0 in the thermodynamic limit rests on a four-point finite-size extrapolation with single training runs, so the significance of the result in its strongest form is not yet fully supported.
major comments (3)
- [Sec. IV B, Fig. 4(b)] The central claim that learned strategies eliminate the volume-law phase for any fixed p>0 is an extrapolation from N=16,32,64,128, with each point apparently coming from a single trained policy and no training-seed error bars. The self-similar pyramid argument rests on one doubling step (2 to 4 pyramids from N=64 to N=128 at p=0.05). The paper itself notes deviations from the proposed (pN)^{-1} law and concedes that training can converge to local minima that eliminate pyramids. Please provide larger system sizes (e.g., N=256) and multiple independent training seeds, or revise the claim to a finite-size evidence statement.
- [Eqs. (6) and (7)] The simplified scaling model is algebraically inconsistent. Equation (6), read as p = 1 - (p/N)n_B, gives n_B = N(1-p)/p and hence Stot/N ~ p/[(1-p)N], not the printed 1/(p/(1-p)N + 1) ~ (1-p)/(pN). If Eq. (6) was instead intended as p = (1-p/N)^{n_B}, then n_B ~ N ln(1/p)/p and Stot/N ~ p/[N ln(1/p)], also not Eq. (7). This error must be corrected before the (pN)^{-1} scaling can be used to support the thermodynamic-limit inference.
- [Sec. IV C, Fig. 5(a)] The partial-information critical threshold qc(p) is also inferred from single trained models per (p,N,q), and the phase boundary in Fig. 5(b) is a guide-to-eye line without an operational definition or uncertainty estimate. Repeated training seeds and a finite-size scaling analysis are needed to establish the bifurcation and the claimed trade-off between q and p.
minor comments (5)
- [Acknowledgments] Typo: 'AKNOWLEDGMENTS' should be 'ACKNOWLEDGMENTS'.
- [Appendix D 6] The text refers to the 'Zotero repository' but the availability statement and reference [85] point to Zenodo; please correct the name.
- [Appendix D 1] The environment description contains a duplicated paragraph describing the step dynamics; please remove the repetition.
- [Eq. (2)] The notation N for both the system size and the normalization factor in Eq. (2) is confusing. Consider using a different symbol for the normalization, e.g., N_norm or S_max.
- [Appendix C 4 / Sec. IV B] The statement that learned strategies 'cannot be replaced by simple human-designed control rules' is supported by single realizations of the pyramid strategy in Fig. 11. Reporting statistics over multiple stochastic realizations would make the claim more robust.
Circularity Check
No load-bearing circularity: central results are direct simulations; the scaling model is an explicit explanatory toy model, and self-citations are background only.
full rationale
I walked the paper's derivation chain and found no step where a claimed prediction is equivalent to its input by construction. The central entanglement results (Sec. IV B, Fig. 4) are direct numerical measurements of a well-defined stabilizer-circuit stochastic process. The RL reward is indeed $r_t=-\langle S_{\rm tot}\rangle/N$ (Eq. 5), i.e., the same order parameter that is later reported; however, the paper's physical claim is not merely that the agent lowers the reward, but that the optimized policy changes the scaling with system size, producing area-law behavior for fixed $p>0$ and above an information threshold $q_c(p)$. That scaling is not encoded in the reward and is measured after training on independent circuits, so the optimization objective does not make the scaling claim circular. The simplified scaling model in Eqs. (6)--(7) is explicitly introduced as a model that 'captures essential features and provides useful intuition' while 'disregarding some underlying physics'; it is a stated-ansatz explanatory construction, not a fit to the numerical data, and the authors themselves note deviations from its $(pN)^{-1}$ prediction due to finite-size effects and local minima in the RL loss landscape. Thus it does not constitute a fitted parameter renamed as a prediction. Self-citations (Refs. [40,45,46]) appear, but they are used for background or for contrast with earlier RL disentangling work, not as the load-bearing justification for the central claims; the random-strategy critical point is checked against the external Ref. [37]. The main caveat is that the thermodynamic-limit statement rests on finite-size extrapolation from $N=64$ to $N=128$ with limited pyramid-doubling evidence and acknowledged deviations; that is an inference-robustness concern, not circularity.
Assumptions & free parameters
assumptions (7)
- standard math Gottesman-Knill theorem: Clifford circuits with classical control are efficiently simulable.
- standard math Clipped gauge relation: entanglement entropy across a cut equals half the number of stabilizer endpoints crossing the cut.
- domain assumption Partial observability is modeled by dropping tableau rows independently with probability 1-q.
- domain assumption PPO training converges to a near-optimal policy for each (p,N,q).
- ad hoc to paper Self-similar pyramid scaling continues beyond N=128.
- ad hoc to paper Simplified model assumptions: independent bond entanglement probabilities p/N and (1-p)/N and steady-state condition Eq. (6).
- domain assumption In the q to 0 limit, targeting Bell-like pairs is sufficient to maintain area-law scaling for p>=pgreedy_c.
Cite this review
Pith. "Pith review of Learning to stabilize nonequilibrium phases of matter with active feedback using partial information." pith.science (2026). https://pith.science/paper/F455LNAR
@misc{pith2026250806612,
author = {Pith},
title = {Pith review of: Learning to stabilize nonequilibrium phases of matter with active feedback using partial information},
year = {2026},
howpublished = {\url{https://pith.science/paper/F455LNAR}},
note = {Machine review of arXiv:2508.06612}
}
read the original abstract
We investigate the role of information in active feedback control of quantum many-body systems using reinforcement learning. Active feedback breaks detailed balance, enabling the engineering of steady states and dynamical phases of matter otherwise inaccessible in equilibrium. We train reinforcement learning agents using partial state information to prevent entanglement spreading in (1+1)-dimensional stabilizer circuits with up to 128 qubits. We find that, above a critical information threshold, learned near-optimal strategies are non-greedy, stochastic, and reduce volume-law entangled steady states to area-law scaling. The agents achieve this by placing a series of bottlenecks that induce pyramidal structures in the long-time spatial entanglement distribution, which effectively split the system and reduce the maximum accessible entanglement. Crucially, learned strategies are inherently out of equilibrium and require real-time active feedback; we find that the learned behavior cannot be replaced by simple human-designed control rules. This work establishes the foundations for classically implemented, information-driven individual control of many interacting quantum degrees of freedom, demonstrating the capabilities of reinforcement learning to stabilize and uncover novel critical properties of many-body nonequilibrium steady states.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Measuring transient times to steady state 13
-
[2]
Enhanced data visualization 14
-
[3]
Critical point determination via Binder cumulant 14
-
[4]
Stability of human-designed pyramid strategies 14
-
[5]
Reinforcement learning methodology 16
Learned strategies with vanishing information 14 D. Reinforcement learning methodology 16
-
[6]
Environment setup and dynamics 16
-
[7]
Training Algorithm 17
-
[8]
Policy Architecture 18
Show all 123 references
-
[9]
Computational Resources 19
-
[10]
Stabilizer circuits 20
Hyperparameters 20 E. Stabilizer circuits 20
-
[11]
Disentangling Clifford gates 23
Simulating Stabilizer circuits 22 F. Disentangling Clifford gates 23
-
[12]
Efficient clipping algorithm 23 a. Performance Considerations 25 Appendix A: Description of supplementary videos The paper is accompanied by twelve supplementary videos illustrating the dynamics of the entanglement pro- file [see Fig.1(b)] for both greedy and learned strategie...
-
[13]
To this end, we simulate 2 11 independent trajec- 14 tories, each initialized in the product state |0⟩⊗N , for various values of the bias parameter p and system size N
Measuring transient times to steady state In this section, we outline the measurement of the tran- sient time required for the circuit to reach the steady state. To this end, we simulate 2 11 independent trajec- 14 tories, each initialized in the product state |0⟩⊗N , for vari...
-
[14]
2 but with improved visual clarity through sep- arate panels
Enhanced data visualization Figure 9 presents the same dataset shown in the main text Fig. 2 but with improved visual clarity through sep- arate panels. This avoids overlapping curves of different strategies and better highlights the distinct scaling be- haviors and critical t...
-
[15]
The fourth- order cumulant is defined as: U = 1 − ⟨x4⟩ 3⟨x2⟩2 (C2) where x = Stot/N is the normalized total entanglement
Critical point determination via Binder cumulant To determine the critical bias for the greedy strategy, we employ Binder cumulant analysis [84]. The fourth- order cumulant is defined as: U = 1 − ⟨x4⟩ 3⟨x2⟩2 (C2) where x = Stot/N is the normalized total entanglement. Figure 10...
-
[16]
Stability of human-designed pyramid strategies In Section IV B, we noted that learned strategies, while appearing simple, are surprisingly difficult to replicate using human-designed approaches. As evidence, we con- sider the failure of a human-designed ”pyramid strategy”: for...
-
[17]
Here, we present numerical evidence supporting this conjecture
Learned strategies with vanishing information In Section IV C, we conjecture that for p > pgreedy c , area-law scaling persists as q → 0. Here, we present numerical evidence supporting this conjecture. The underlying intuition is that, as the system size grows, the gate densit...
-
[18]
II), we implement a custom RL environment
Environment setup and dynamics To learn and investigate entanglement-reducing strate- gies in Clifford circuits (Sec. II), we implement a custom RL environment. The RL pipeline, outlined in Sec. IV A and Fig. 3, trains an agent to learn a policy π(a|s) that selects positions a...
-
[19]
Policy gradient methods directly learn a parameterized policy πθ by updating the parameters θ to maximize expected rewards
T raining Algorithm To train the RL agent, we use the Proximal Policy Op- timization (PPO) algorithm [63], a state-of-the-art policy gradient method. Policy gradient methods directly learn a parameterized policy πθ by updating the parameters θ to maximize expected rewards. The...
-
[20]
Specifically, we employ the en- coder component of the transformer architecture illus- trated in Fig
Policy Architecture Both the actor and critic are modeled with a transformer-based architecture implemented using the Flax Python library. Specifically, we employ the en- coder component of the transformer architecture illus- trated in Fig. 13. The encoder architecture stacks ...
-
[21]
T raining Process We employ a warmup cosine decay learning rate sched- ule which improves training stability, facilitates conver- gence, and enables soft restarts from checkpoints – partic- ularly at large system sizes ( N = 64, 128) where training required multiple checkpoint...
-
[22]
However, the dominant bottleneck is the transient time required for the physical dynamics to reach the steady state, which scales as Ttr ∝ N 3
Computational Resources All training are conducted on a high-performance com- puting system with the following specifications: • Intel Xeon Platinum 8360Y CPUs (36 cores at 2.40GHz, dual CPU per node) • 72 execution nodes with 1 TB RAM and 4 Nvidia A100-80GB GPUs each Resource...
-
[23]
While common defaults were used when possible, some fine- tuning was required for each model
Hyperparameters Table I summarizes the hyperparameters used for training across system sizes N ∈ 16, 32, 64, 128. While common defaults were used when possible, some fine- tuning was required for each model. The complete list of hyperparameter configurations is available in th...
-
[24]
The cen- tral component of the algorithm is the so-called tableau that represents the n stabilizer and the n destabilizer gen- erators
Simulating Stabilizer circuits First introduced by Aaronson and Gottesman [54], the Tableau algorithm represents an efficient method to sim- ulate stabilizer circuits on classical computers. The cen- tral component of the algorithm is the so-called tableau that represents the ...
-
[25]
This is done at each time step of the dynamics, and an efficient clipping algorithm is essential to keep computa- tion time manageable
Efficient clipping algorithm Given a stabilizer state |ψ⟩, mapping it to disentan- gling gates requires the tableau to be in clipped gauge. This is done at each time step of the dynamics, and an efficient clipping algorithm is essential to keep computa- tion time manageable. A...
-
[26]
The row must not have been processed
-
[27]
It must have an endpoint on the current column (left or right, depending on direction)
-
[28]
Once the pivot is found, it is used to eliminate all other rows with an endpoint on the same column
Among these, the row with the smallest support is chosen. Once the pivot is found, it is used to eliminate all other rows with an endpoint on the same column. This pro- ceeds as follows:
-
[29]
Identify the pivot row
-
[30]
Create a boolean mask to select rows that need updating
-
[31]
This parallelized approach significantly reduces compu- tation time
Apply a single bitwise XOR operation to update all selected rows in parallel. This parallelized approach significantly reduces compu- tation time. Finally, the clipped gauge also enables direct computa- tion of entanglement. The number of stabilizers crossing a given cut deter...
-
[32]
Landauer, IBM J
R. Landauer, IBM J. Res. Dev. 5, 183 (1961)
1961
-
[33]
Salath´ e, P
Y. Salath´ e, P. Kurpiers, T. Karg, C. Lang, C. K. An- dersen, A. Akin, S. Krinner, C. Eichler, and A. Wallraff, Phys. Rev. Appl. 9, 034011 (2018)
2018
-
[34]
Sakaguchi, S
A. Sakaguchi, S. Konno, F. Hanamura, W. Asavanant, H. Yamazaki, A. Youssefi, K. Takase, J.-i. Yoshikawa, N. C. Menicucci, K. Ping, K. Miyata, H. Ogawa, P. Marek, R. Filip, and A. Furusawa, Nat. Commun. 14, 3817 (2023)
2023
-
[35]
Iqbal, N
M. Iqbal, N. Tantivasadakarn, T. M. Gatterman, J. A. Gerber, K. Gilmore, D. Gresh, A. Hankin, N. Hewitt, C. V. Horst, M. Matheny, T. Mengle, B. Neyenhuis, M. Foss-Feig, A. Vishwanath, R. Verresen, and H. Dreyer, Commun. Phys. 7, 205 (2024)
2024
-
[36]
Buhrman, M
H. Buhrman, M. Folkertsma, B. Loff, and N. M. P. Neu- mann, Quantum 8, 1552 (2024)
2024
-
[37]
Wendin, Rep
G. Wendin, Rep. Prog. Phys. 80, 106001 (2017)
2017
- [38]
-
[39]
D. J. Gross, P. Zoller, and A. Sevrin, eds., The Physics of Quantum Information: Proceedings of the 28th Solvay Conference on Physics, World Scientific Series on Quan- tum Science and Technology (World Scientific, Singa- pore, 2023) brussels, Belgium, 19–21 May 2022
2023
-
[40]
J. D. Biamonte, M. E. S. Morales, and D. E. Koh, Phys. Rev. A 101, 012349 (2020)
2020
-
[41]
A. J. Daley, I. Bloch, C. Kokail, S. Flannigan, N. Pearson, M. Troyer, and P. Zoller, Nature 607, 667 (2022)
2022
-
[42]
G. E. Fux, E. Tirrito, M. Dalmonte, and R. Fazio, Phys. Rev. Res. 6, L042030 (2024)
2024
-
[43]
Q. Zhao, Y. Zhou, and A. M. Childs, Nat. Phys. 10.1038/s41567-025-02945-2 (2025)
2025 doi
-
[44]
Aolita, F
L. Aolita, F. de Melo, and L. Davidovich, Rep. Prog. Phys. 78, 042001 (2015)
2015
-
[45]
Bauer, S
B. Bauer, S. Bravyi, M. Motta, and G. K. Chan, Chem. Rev. 120, 12685 (2020)
2020
-
[46]
D ´ ıez-Valle, D
P. D ´ ıez-Valle, D. Porras, and J. J. Garc ´ ıa-Ripoll, Phys. Rev. A 104, 062426 (2021)
2021
-
[47]
Niedermeier, J
M. Niedermeier, J. L. Lado, and C. Flindt, Phys. Rev. Res. 6, 033325 (2024)
2024
-
[48]
A. C. Nakhl, T. Quella, and M. Usman, Phys. Rev. A 109, 032413 (2024)
2024
-
[49]
G. J. Mooney, G. A. L. White, C. D. Hill, and L. C. L. Hollenberg, npj Quantum Inf. 7, 27 (2021)
2021
-
[50]
Zhang, T
X.-M. Zhang, T. Li, and X. Yuan, Phys. Rev. Lett. 129, 230504 (2022)
2022
-
[51]
G.-Y. Zhu, N. Tantivasadakarn, A. Vishwanath, S. Trebst, and R. Verresen, Phys. Rev. Lett. 131, 200201 (2023)
2023
-
[52]
E. H. Chen, G.-Y. Zhu, R. Verresen, A. Seif, E. Ba¨ umer, D. Layden, N. Tantivasadakarn, G. Zhu, S. Shel- don, A. Vishwanath, S. Trebst, and A. Kandala, arXiv:2309.02863 (2023)
2023 arXiv
-
[53]
Islam, R
R. Islam, R. Ma, P. M. Preiss, M. E. Tai, A. Lukin, M. Rispoli, and M. Greiner, Nature 528, 77 (2015)
2015
-
[54]
Vermersch, A
B. Vermersch, A. Elben, M. Dalmonte, J. I. Cirac, and P. Zoller, Phys. Rev. A 97, 023604 (2018)
2018
-
[55]
Brydges, M
M. Brydges, M. Martin, D. Eldredge, J. H. Wesenberg, D. H. J. K. Dahlberg, C. A. Muschik, M. B. Plenio, A. M. Rey, and C. Monroe, Science 364, 260 (2019)
2019
-
[56]
Vasseur and J
R. Vasseur and J. E. Moore, J. Stat. Mech. 2016, 064010 (2016)
2016
-
[57]
Nahum, S
A. Nahum, S. Vijay, and J. Haah, Phys. Rev. X8, 021014 (2018)
2018
-
[58]
Swingle, Nat
B. Swingle, Nat. Phys. 14, 988 (2018)
2018
-
[59]
Y. Li, X. Chen, and M. P. A. Fisher, Phys. Rev. B 100, 134306 (2019)
2019
-
[60]
M. P. Fisher, V. Khemani, A. Nahum, and S. Vijay, Annu. Rev. Condens. Matter Phys. 14, 335 (2023)
2023
-
[61]
Y. Li, X. Chen, and M. P. A. Fisher, Phys. Rev. B 98, 205136 (2018)
2018
-
[62]
Skinner, J
B. Skinner, J. Ruhman, and A. Nahum, Phys. Rev. X 9, 031009 (2019)
2019
-
[63]
A. Chan, R. M. Nandkishore, M. Pretko, and G. Smith, Phys. Rev. B 99, 224307 (2019)
2019
-
[64]
C. Noel, P. Niroula, D. Zhu, A. Risinger, L. Egan, D. Biswas, M. Cetina, A. V. Gorshkov, M. J. Gullans, D. A. Huse, and C. Monroe, Nat. Phys. 18, 760 (2022)
2022
-
[65]
J. C. Hoke, M. Ippoliti, E. Rosenberg, D. Abanin, R. Acharya, T. I. Andersen, M. Ansmann, F. Arute, K. Arya, A. Asfaw, ..., and P. Roushan, Nature 622, 481 (2023)
2023
-
[66]
J. Koh, S. Sun, M. Motta, A. J. Minnich, and ..., Nat. Phys. 19, 1314 (2023)
2023
-
[67]
Nahum, J
A. Nahum, J. Ruhman, S. Vijay, and J. Haah, Phys. Rev. X 7, 031016 (2017)
2017
-
[68]
Morral-Yepes, A
R. Morral-Yepes, A. Smith, S. Sondhi, and F. Pollmann, PRX Quantum 5, 010309 (2024)
2024
-
[69]
Morral-Yepes, M
R. Morral-Yepes, M. Langer, A. Gammon-Smith, B. Kraus, and F. Pollmann, arXiv:2507.05055 (2025)
2025 arXiv
-
[70]
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. (MIT Press, 2018)
2018
-
[71]
Bukov, A
M. Bukov, A. G. R. Day, D. Sels, P. Weinberg, A. Polkovnikov, and P. W. Claeys, Phys. Rev. X 8, 031086 (2018)
2018
-
[72]
F¨ osel, P
T. F¨ osel, P. Tighineanu, T. Weiss, and F. Marquardt, Phys. Rev. X 8, 031084 (2018)
2018
-
[73]
M. Y. Niu, S. Boixo, V. N. Smelyanskiy, and H. Neven, npj Quantum Inf. 5, 33 (2019). 26
2019
-
[74]
M. M. Wauters, E. Panizon, G. B. Mbeng, and G. E. Santoro, Phys. Rev. Res. 2, 033446 (2020)
2020
-
[75]
S.-F. Guo, F. Chen, Q. Liu, M. Xue, J.-J. Chen, J.-H. Cao, T.-W. Mao, M. K. Tey, and L. You, Phys. Rev. Lett. 126, 060401 (2021)
2021
-
[76]
Metz and M
F. Metz and M. Bukov, Nat. Mach. Intell. 5, 780 (2023)
2023
-
[77]
Tashev, S
P. Tashev, S. Petrov, F. Metz, and M. Bukov, arXiv preprint arXiv:2406.07884 (2024), arXiv:2406.07884 [quant-ph]
2024 arXiv
-
[78]
J. Olle, R. Zen, M. Puviani, and F. Marquardt, npj Quan- tum Inf. 10, 126 (2024)
2024
-
[79]
V. B. Sørdal and J. Bergli, Phys. Rev. A 100, 042314 (2019)
2019
-
[80]
S. R. P, arXiv:2305.06177 (2023)
2023 arXiv
-
[81]
P. A. Erdman, R. Czupryniak, B. Bhandari, A. N. Jor- dan, F. No´ e, J. Eisert, and G. Guarnieri, Quantum Sci. Technol. 10, 025047 (2025)
2025
-
[82]
Gottesman, Phys
D. Gottesman, Phys. Rev. A 54, 1862 (1996)
1996
-
[83]
Gottesman, Phys
D. Gottesman, Phys. Rev. A 57, 127–137 (1998)
1998
- [84]
-
[85]
Aaronson and D
S. Aaronson and D. Gottesman, Phys. Rev. A 70, 052328 (2004)
2004
-
[86]
S. Sang, Z. Li, T. H. Hsieh, and B. Yoshida, PRX Quan- tum 4, 040332 (2023)
2023
-
[87]
Richter, O
J. Richter, O. Lunt, and A. Pal, Phys. Rev. Res. 5, L012031 (2023)
2023
-
[88]
O. Lunt, M. Szyniszewski, and A. Pal, Phys. Rev. B 104, 155111 (2021)
2021
-
[89]
Makki, N
N. Makki, N. Lang, and H. P. B¨ uchler, Phys. Rev. Res. 6, 013278 (2024)
2024
-
[90]
J. Z. Zhuang and colleagues, Phys. Rev. Research 5, L042043 (2023)
2023
-
[91]
The critical value pc is smaller than 0 .5 due to an asym- metry in the gates. The entanglement-reducing strategy exclusively places gates that reduce entanglement, while the random strategy samples gates uniformly from the Clifford group C2 – most of which increase entangleme...
-
[92]
Nandkishore and D
R. Nandkishore and D. A. Huse, Annu. Rev. Condens. Matter Phys. 6, 15 (2015)
2015
-
[93]
The RL environment is conceptually distinct from the notion of environment in open systems in physics
-
[94]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, arXiv:1707.06347 (2017)
2017 arXiv
-
[95]
Freeman, E
D. Freeman, E. Frey, A. Tsividis, S. Yazdani, J. Knight, J. S. Park, S. Levine, and H. Michalewski, Brax - a dif- ferentiable physics engine for large scale rigid body sim- ulation (2021)
2021
-
[96]
Konda and J
V. Konda and J. Tsitsiklis, in Adv. Neural Inf. Process. Syst., Vol. 12 (MIT Press, 1999)
1999
-
[97]
Bradbury, R
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. Van- derPlas, S. Wanderman-Milne, and Q. Zhang, JAX: com- posable transformations of Python+NumPy programs (2018)
2018
-
[98]
Cramer, M
M. Cramer, M. B. Plenio, S. T. Flammia, R. Somma, D. Gross, S. D. Bartlett, O. Landon-Cardinal, D. Poulin, and Y. Liu, Nat. Commun. 1, 149 (2010)
2010
-
[99]
B. P. Lanyon, C. Maier, M. Holz¨ apfel, T. Baumgratz, C. Hempel, P. Jurˇ cevi´ c, I. Dhand, A. S. Buyskikh, A. J. Daley, M. Cramer, M. B. Plenio, R. Blatt, and C. F. Roos, Nat. Phys. 13, 1158 (2017)
2017
-
[100]
Rocchetto, S
A. Rocchetto, S. Aaronson, S. Severini, G. Carvacho, D. Poderini, I. Agresti, M. Bentivegna, and F. Sciarrino, Sci. Adv. 5, eaau1946 (2019)
2019
-
[101]
A Bell-like pair is a two qubits state that is maximally entangled and locally equivalent (up to single-qubit Clif- ford operation) to a Bell pair
-
[102]
P¨ utz, S
M. P¨ utz, S. J. Garratt, H. Nishimori, S. Trebst, and G. Zhu, arXiv:2504.12385 (2025)
2025 arXiv
-
[103]
Reuer, J
K. Reuer, J. Landgraf, T. F¨ osel, J. O’Sullivan, L. Beltr´ an, A. Akin, G. Norris, A. Remm, M. Kerschbaum, J. Besse, F. Marquardt, A. Wallraff, and C. Eichler, Nat. Com- mun. 14, 7138 (2023)
2023
-
[104]
D. T. R. Nagy, C. Czab´ an, B. Bak´ o, P. H´ aga, Z. Kallus, and Z. Zimbor´ as, arXiv:2410.18284 (2024)
2024 arXiv
-
[105]
Wen, Rev
X.-G. Wen, Rev. Mod. Phys. 89, 041004 (2017)
2017
-
[106]
Senthil, Annu
T. Senthil, Annu. Rev. Condens. Matter Phys. 6, 299 (2015)
2015
-
[107]
Y. Zhou, Q. Zhao, X. Yuan, and X. Ma, NPJ Quantum Inf. 5, 83 (2019)
2019
-
[108]
B. M. Terhal, Rev. Mod. Phys. 87, 307 (2015)
2015
- [109]
-
[110]
F. J. Cao and M. Feito, Phys. Rev. E 79, 041118 (2009)
2009
-
[111]
J. V. Koski, V. F. Maisi, T. Sagawa, and J. P. Pekola, Phys. Rev. Lett. 113, 030601 (2014)
2014
-
[112]
Sagawa and M
T. Sagawa and M. Ueda, Phys. Rev. E 85, 021104 (2012)
2012
-
[113]
Prech and P
K. Prech and P. P. Potts, Phys. Rev. Lett. 133, 140401 (2024)
2024
-
[114]
J. M. Horowitz and M. Esposito, Phys. Rev. X 4, 031015 (2014)
2014
-
[115]
Binder, Phys
K. Binder, Phys. Rev. Let. 47, 693 (1981)
1981
-
[116]
Cemin, M
G. Cemin, M. Schmitt, and M. Bukov, Data, code and models., https://doi.org/10.5281/zenodo. 16672380 (2025), available upon reasonable request
2025 doi
-
[117]
Gottesman, Surviving as a quantum computer in a classical world (2024), textbook manuscript preprint
D. Gottesman, Surviving as a quantum computer in a classical world (2024), textbook manuscript preprint
2024
- [118]
-
[119]
M. A. Nielsen and I. L. Chuang, Quantum Computa- tion and Quantum Information: 10th Anniversary Edi- tion (Cambridge University Press, Cambridge, 2010)
2010
-
[120]
The normalizer of a subset S in a group G is the set of el- ements of G that leave the set S fixed under conjugation, i.e., NG(S) := {g ∈ G|gS = Sg} = {g ∈ G|gSg−1 = S}
-
[121]
Koenig and J
R. Koenig and J. A. Smolin, J. Math. Phys. 55, 122202 (2014)
2014
-
[122]
Ozols, Clifford Group, Essay, University of Waterloo (2008)
M. Ozols, Clifford Group, Essay, University of Waterloo (2008)
2008
-
[123]
Iwould be in S, but −I does not stabilize anything
If S ∈ S, then S can only have a phase of ±1, not ±i; for in the latter case S2 = −I . . . Iwould be in S, but −I does not stabilize anything
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.