REVIEW 4 major objections 5 minor 89 references
Evolutionary reinforcement learning of dynamical large deviations
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Evolutionary search over a reference model's rates reproduces the large-deviation rate function for rare dynamical events, tightly enough that the upper bound alone plots the correct curve.
desk verdict A readable proof-of-principle that evolutionary learning can fit tight variational bounds for dynamical large deviations; honest about its limits, worth refereeing, but not self-certifying beyond known answers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the change-of-dynamics reweighting identity that turns a reference model into a rate-function estimate. With original rates $W_{xy}$, reference rates $\widetilde W_{xy}$, escape rates $R_x$ and $\widetilde R_x$, and jump time $\tilde\Delta t_n$ under the reference dynamics, the per-jump log-ratio $q_{xy}=\ln(W_{xy}/\widetilde W_{xy})-\tilde\Delta t_n(R_x-\widetilde R_x)$ is accumulated along one long reference trajectory; its negative time average, $J_0=-\langle q\rangle_{\tilde a_0}^{\mathrm{ref}}$, is an upper bound on $J(\tilde a_0)$ by Jensen's inequality. The evolutionary machinery that minimizes this bound consists of multiplicative rate mutations $\hat W_{xy}=e^{\epsilon(\eta_{xy}-1/2)}\widetilde W_{xy}$ (or Gaussian shifts of neural-network weights) followed by two selection rules: 'a-evolution' accepts mutations that bring the observable closer to a target value, and 'J-evolution' accepts mutations that lower $J_0$ while keeping the observable within a tolerance. For lattice models the reference rates are parameterized by a single-layer network of spin filters—hidden nodes that detect local patterns of up/down spins—whose output $f_x$ enters through $\widetilde W_{xy}=W_{xy}e^{w_0}e^{f_y-f_x}$.
What would settle it
Run the same evolutionary protocol on a model with an exactly known $J(a)$ whose driven process cannot be represented by the chosen filter set—for example, a lattice activity model in a regime with a dynamical phase transition—and inspect the plateau value of $J_0$ after $J$-evolution has stopped changing it; if at some target $a$ that plateau lies visibly above $J(a)$, then the paper's claim that the bound alone suffices for plotting does not hold for that case, and the method needs the correction term or a richer ansatz.
Extended reading notes
Core claim
The central claim is that evolutionary reinforcement learning can find a reference model whose typical dynamics is effectively the driven (conditioned) dynamics of the original model, so that the change-of-dynamics bound $J_0(\tilde a_0)$ computed from a single long reference-model trajectory is a tight estimate of the exact rate function $J(\tilde a_0)$. The paper demonstrates this for three models: entropy production in a 4-state model, activity in a 15-site Fredrickson-Andersen model, and activity in a 100-site Fredrickson-Andersen model. In each case the evolutionary bound is compared with an exact answer (diagonalization or matrix product states) and lies on or very close to it, with the VARD correction term $J_1$ small enough to be ignored for the purpose of plotting. The authors do not claim the bound is always exact; they claim that the evolutionary search makes it tight in practice for the cases tested, and that the approach requires neither physical insight nor explicit use of large-deviation theory.
Load-bearing premise
The load-bearing assumption is that the evolutionary search actually finds a reference model whose typical behavior is close to the rare conditioned behavior of the original model, and the paper gives no diagnostic to distinguish a converged tight bound from a converged loose one.
Editorial extensions
If this is right
- For any target value $\tilde a_0$, a converged reference model yields a certified upper bound on $J(\tilde a_0)$ from a single trajectory, with statistical error that shrinks like $1/\sqrt{N}$ in the number of events.
- In the three models tested, the bound is tight enough that the second-stage correction term of the VARD method can be skipped for plotting purposes, shortening the calculation to one evolutionary pass.
- Because the search operates directly on rates or network weights, the same loop can be pointed at any path-extensive observable without re-deriving theory for that observable.
- The method's accuracy on the 100-site Fredrickson-Andersen model, where exact answers require matrix-product-state simulation, shows that the evolutionary bound can approach state-of-the-art accuracy with modest desktop CPU time.
Reading between the lines
- A stress test the paper does not run would be a model with a dynamical phase transition, where the rate function has a non-convex or kinked segment; a single neural-network ansatz may then plateau at a bound visibly above $J(a)$, signalling the need for a deeper network or a population of reference models.
- The selection rules are a form of genetic algorithm, so adding mutation-rate adaptation, elitism, or an archive of reference models could provide a stopping criterion that the paper currently lacks.
- Because the bound comes from one long typical trajectory, the scheme could run as a passive diagnostic inside an existing simulation: accumulating $a$ and $J_0$ during production runs would yield rate-function estimates as a by-product.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an evolutionary reinforcement-learning algorithm to compute dynamical large-deviation rate functions. A reference stochastic model, parameterized either by its rates (for small state spaces) or by a neural network (for larger state spaces), is evolved so that its typical dynamics matches the rare dynamics of the original model conditioned on a value of a path-extensive observable. The bound J0 (Eq. 3) is minimized by a-evolution and J-evolution steps. The paper demonstrates on three models (a 4-state entropy-production model, a 15-site Fredrickson-Andersen model, and a 100-site FA model) that the evolved upper bound lies close to, or effectively on, the exact rate function obtained by diagonalization or matrix-product-state methods. The correction term J1 is not computed, and the paper argues that it is small enough for plotting purposes based on the observed agreement with exact results.
Significance. If the method proves robust, it would offer a conceptually simple, model-agnostic route to computing large-deviation rate functions without physical insight or rare-event sampling, complementing existing techniques such as cloning and transition-path sampling. The introduction of neural-network reference-model ansaetze and evolutionary optimization into dynamical large deviations is a useful contribution. The paper is honest in comparing with existing bounds and in validating against exact/MPS results for all three models. However, the central claim that the method 'allows the calculation' of the rate function from the bound alone is not supported by an internal diagnostic; the tightness of the bound is only established a posteriori. This limits the paper to a proof-of-principle demonstration unless the missing certification is supplied.
major comments (4)
- [Sections III and V; Fig. 5(c)] The claim that the bound J0 alone suffices for plotting the rate function rests on the uncomputed correction term J1 (Eq. A17). The paper states in Section III that 'We do not address here the calculation of the correction term' and justifies the smallness of J1 solely by a posteriori comparison with exact/MPS results. No internal diagnostic is provided to certify that the evolved reference model is close to the driven model, and no stopping rule is given to distinguish a converged tight bound from a converged loose bound. In fact, Fig. 5(c) shows an evolutionary trajectory (gray) that has reached a plateau but lies visibly above the exact answer, demonstrating that a plateau is not a certificate of tightness. To support the central claim for models without known answers, the paper should compute J1 (or estimate the Jensen gap) for at least one example, or explicitly restrict the claim to bounds whose tightness must be verified by external means.
- [Section IV.A, Eq. (11)] The text states that the regularized a-evolution requires that 'the new bound must be not more than a value mu = 0.2 larger than the current bound,' but the displayed inequality is |a_hat - a*| < |a - a*| and a_hat < a + mu, where a is the observable and not the bound. The discussion of mu (e.g., 'the bound must be allowed to increase in size') indicates that the regularization is meant to act on J0, not on a. As written, Eq. (11) is internally inconsistent with the text and prevents faithful reproduction of the algorithm. Please correct the equation and clarify whether the implemented condition was J_hat0 < J0 + mu.
- [Section IV.B and figures] For the L=100 FA model, the paper acknowledges that the evolved bound is 'inexact, but numerically close' to the MPS result, and Fig. 5(b) shows a visible gap between the K=5 bound and the exact curve. The conclusion nonetheless states that for all three models 'the discrepancy between bound and exact answer ... is so small that for the purposes of plotting the rate function no correction is required.' Please quantify the maximum deviation |J0 - J| for each model and state whether this deviation is within symbol sizes or within a stated tolerance. The absence of plotted error bars (despite the assertion that they are smaller than symbol sizes) makes this assessment impossible for the reader.
- [Section III, IV.A, IV.B (evolutionary protocol)] The evolutionary protocol is a monotone hill-climber with several hand-tuned parameters (epsilon, sigma, delta, mu, Nev, N). No repeated runs with identical target values are reported, so the reproducibility and variance of the final bounds are unknown. For a numerical method, it is essential to show that multiple independent seeds lead to statistically consistent results for a fixed target a*. Such a test would also provide a direct measure of the ruggedness of the optimization landscape and would complement the plateau-based stopping criterion, which, as noted above, is not a sufficient diagnostic for tightness.
minor comments (5)
- [Appendix A, Eq. (A13)] The definition of q(ω) appears to have an erroneous minus sign: the log-likelihood ratio should be q(ω) = T^{-1} \sum q_{x_n x_{n+1}}, not -T^{-1} \sum q_{x_n x_{n+1}}, given the definition of q_{xy} in Eq. (A14). The final expression for J0 is consistent with the main text, so this is likely a typo, but it should be corrected to avoid confusing readers.
- [Throughout] The symbol '&' is used where a mathematical comparison is intended, e.g., 'for K & 4' in Sections IV.A and IV.B. This should read 'K ≥ 4' or is otherwise ambiguous.
- [Section IV.A, Eq. (10)] The statement 'Error bars associated with the bound scale as 1/\sqrt{N}' is made, but no error bars appear in any figure. Please include representative error bars (or an explicit statement covering all plotted points) so that the claim that bounds are 'numerically close' can be quantitatively assessed.
- [Figure 3(c)] The space-time plots would benefit from a clearer caption: the vertical axis is the lattice site, the horizontal axis time, and the numbers on the left are the typical activity values; please state this explicitly in the caption or in the text.
- [Appendix B, Eq. (B2)] The derivation of the CMP bound replaces the fluctuating jump time \tilde{\Delta t}_n with its mean 1/\tilde{R}_x in Eq. (3). This replacement is not fully justified for an unbounded sum and deserves a more careful statement about the limit in which this approximation becomes exact.
Circularity Check
No circularity: the variational bound is derived self-contained and validated against external exact/MPS results.
full rationale
The paper's central quantity J0 is an upper bound derived in Appendix A from the Radon-Nikodym reweighting (Eqs. A11-A18); Jensen's inequality gives J <= J0. The evolutionary algorithm minimizes this bound for each target a* without fitting to the exact rate function. The reported points are computed from simulated reference-model trajectories, and their agreement with the exact rate function is checked against independent matrix-diagonalization and MPS benchmarks (Figs. 1, 3-5). The only notable self-citation is to Ref. [51] for the VARD framework and for the existence of a correction term J1, but the paper explicitly does not compute J1 and does not use Ref. [51] to certify tightness; tightness is an empirical finding from external comparisons. The lack of a convergence guarantee for the evolutionary search and the uncomputed correction term are correctness/robustness limitations, not circular reasoning. No fitted parameter is relabeled as a prediction, and no equation reduces to its own input by construction.
Assumptions & free parameters
free parameters (10)
- Reference-model rates of the 4-state model (12 rates) =
not tabulated; resulting models shown in Fig. 1(a)
- Neural-network weights for L=15 FA model (w0, w1, {w_k^alpha}, 2K+1 parameters) =
shown in Fig. 3(b)
- Neural-network weights for L=100 FA model (w0 and 2^K weights) =
not tabulated
- Filter order K =
K=7 (L=15), K=5 (L=100)
- Evolutionary rate epsilon =
0.1 (a-evolution), 0.05 (J-evolution)
- Tolerance delta for J-evolution =
0.1 (4-state), 0.02 (FA)
- Regularization mu =
0.2
- Trajectory length N =
10^4 (4-state), 10^5 (L=15 FA), 2x10^5 (L=100 FA)
- Number of final J-evolution trajectories Nev =
10^5 (4-state), 3x10^4 (FA)
- Mutation step sigma for network weights =
0.01
assumptions (6)
- domain assumption Large-deviation scaling rho_T(A) ~ e^{-T J(a)} holds for the models considered.
- standard math The continuous-time Monte Carlo process is correctly described by exponential jump times and rate-proportional jump probabilities (Gillespie).
- standard math The bound J(a0) <= J0(a0) follows from Jensen's inequality applied to the reweighting identity.
- domain assumption The VARD correction term J1 is assumed small for the optimized reference models without computing it.
- ad hoc to paper The driven model for the FA models is well approximated by the neural-network ansatz with filters of order K (K>=4).
- ad hoc to paper The evolutionary search does not get trapped in a poor local minimum that yields a loose bound.
Cite this review
Pith. "Pith review of Evolutionary reinforcement learning of dynamical large deviations." pith.science (2026). https://pith.science/paper/RGZSS5XR
@misc{pith2026190900835,
author = {Pith},
title = {Pith review of: Evolutionary reinforcement learning of dynamical large deviations},
year = {2026},
howpublished = {\url{https://pith.science/paper/RGZSS5XR}},
note = {Machine review of arXiv:1909.00835}
}
read the original abstract
We show how to calculate the likelihood of dynamical large deviations using evolutionary reinforcement learning. An agent, a stochastic model, propagates a continuous-time Monte Carlo trajectory and receives a reward conditioned upon the values of certain path-extensive quantities. Evolution produces progressively fitter agents, eventually allowing the calculation of a piece of a large-deviation rate function for a particular model and path-extensive quantity. For models with small state spaces the evolutionary process acts directly on rates, and for models with large state spaces the process acts on the weights of a neural network that parameterizes the model's rates. This approach shows how path-extensive physics problems can be considered within a framework widely used in machine learning.
Figures
Reference graph
Works this paper leans on
-
[1]
Behler and M
J. Behler and M. Parrinello, Phys. Rev. Lett. 98, 146401 (2007)
2007
-
[2]
Mills, M
K. Mills, M. Spanner, and I. Tamblyn, Physical Review A 96, 042113 (2017)
2017
-
[3]
A. L. Ferguson and J. Hachmann, Molecular Systems De- sign & Engineering (2018)
2018
-
[4]
Artrith, A
N. Artrith, A. Urban, and G. Ceder, The Journal of Chemical Physics 148, 241711 (2018)
2018
-
[5]
Singraber, T
A. Singraber, T. Morawietz, J. Behler, and C. Dellago, Journal of Physics: Condensed Matter 30, 254005 (2018)
2018
-
[6]
Desgranges and J
C. Desgranges and J. Delhommelle, The Journal of Chemical Physics 149, 044118 (2018)
2018
-
[7]
Thurston and A
B. Thurston and A. Ferguson, Molecular Simulation , 1 (2018)
2018
-
[8]
Singraber, J
A. Singraber, J. Behler, and C. Dellago, Journal of Chemical theory and computation 15, 1827 (2019)
2019
Show all 89 references
-
[9]
Han et al., arXiv preprint arXiv:1611.07422 (2016)
J. Han et al., arXiv preprint arXiv:1611.07422 (2016)
2016 arXiv
-
[10]
K. T. Sch¨ utt, F. Arbabzadah, S. Chmiela, K. R. M¨ uller, and A. Tkatchenko, Nature Communications 8, 13890 (2017)
2017
-
[11]
K. Yao, J. E. Herr, D. Toth, R. Mckintyre, and J. Parkhill, Chem. Sci. 9, 2261 (2018)
2018
-
[12]
K. T. Sch¨ utt, H. E. Sauceda, P.-J. Kindermans, A. Tkatchenko, and K.-R. Mller, The Journal of Chem- ical Physics 148, 241722 (2018)
2018
-
[13]
Carrasquilla and R
J. Carrasquilla and R. G. Melko, Nature Physics 13, 431 (2017), 1605.01735
2017 arXiv
-
[14]
Portman and I
N. Portman and I. Tamblyn, Journal of Computational Physics 350, 871 (2017)
2017
-
[15]
B. S. Rem, N. K¨ aming, M. Tarnowski, L. Asteria, N. Fl¨ aschner, C. Becker, K. Sengstock, and C. Weiten- berg, Nature Physics (2019), 10.1038/s41567-019-0554-0
2019 doi
-
[16]
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction (2018)
2018
-
[17]
T. I. Ahamed, V. S. Borkar, and S. Juneja, Operations Research 54, 489 (2006)
2006
-
[18]
A. Basu, T. Bhattacharyya, and V. S. Borkar, Mathe- matics of operations research 33, 880 (2008)
2008
-
[19]
V. S. Borkar, in Proceedings of the 19th International Symposium on Mathematical Theory of Networks and 11 Systems–MTNS, Vol. 5 (2010)
2010
-
[20]
Borkar, S
V. Borkar, S. Juneja, A. Kherani, et al., Communications in Information & Systems 3, 259 (2003)
2003
-
[21]
Chetrite and H
R. Chetrite and H. Touchette, Journal of Statistical Me- chanics: Theory and Experiment 2015, P12001 (2015)
2015
-
[22]
H. J. Kappen and H. C. Ruiz, Journal of Statistical Physics 162, 1244 (2016)
2016
-
[23]
Nemoto, R
T. Nemoto, R. L. Jack, and V. Lecomte, Physical Review Letters 118, 115702 (2017)
2017
- [24]
-
[25]
T. A. Bojesen, Phys. Rev. E 98, 063303 (2018)
2018
-
[26]
C. J. Watkins and P. Dayan, Machine learning 8, 279 (1992)
1992
-
[27]
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, arXiv preprint arXiv:1312.5602 (2013)
2013 arXiv
-
[28]
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Ve- ness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al., Nature 518, 529 (2015)
2015
-
[29]
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, Journal of Artificial Intelligence Research 47, 253 (2013)
2013
-
[30]
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, in Interna- tional conference on machine learning (2016) pp. 1928– 1937
2016
-
[31]
Tassa, Y
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al. , arXiv preprint arXiv:1801.00690 (2018)
2018 arXiv
-
[32]
Todorov, T
E. Todorov, T. Erez, and Y. Tassa, in Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ Interna- tional Conference on (IEEE, 2012) pp. 5026–5033
2012
-
[33]
M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming (John Wiley & Sons, 2014)
2014
-
[34]
Asperti, D
A. Asperti, D. Cortesi, and F. Sovrano, arXiv preprint arXiv:1804.08685 (2018)
2018 arXiv
-
[35]
Riedmiller, in European Conference on Machine Learning (Springer, 2005) pp
M. Riedmiller, in European Conference on Machine Learning (Springer, 2005) pp. 317–328
2005
-
[36]
Riedmiller, T
M. Riedmiller, T. Gabel, R. Hafner, and S. Lange, Au- tonomous Robots 27, 55 (2009)
2009
-
[37]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, arXiv preprint arXiv:1707.06347 (2017)
2017 arXiv
-
[38]
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. O. Stanley, and J. Clune, arXiv preprint arXiv:1712.06567 (2017)
2017 arXiv
-
[39]
Brockman, V
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, arXiv preprint arXiv:1606.01540 (2016)
2016 arXiv
-
[40]
Kempka, M
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Ja´ skowski, inComputational Intelligence and Games (CIG), 2016 IEEE Conference on (IEEE, 2016) pp. 1–8
2016
-
[41]
Wydmuch, M
M. Wydmuch, M. Kempka, and W. Ja´ skowski, arXiv preprint arXiv:1809.03470 (2018)
2018 arXiv
-
[42]
Silver, A
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al., nature 529, 484 (2016)
2016
-
[43]
Silver, J
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al., Nature 550, 354 (2017)
2017
-
[44]
Touchette, Physics Reports 478, 1 (2009)
H. Touchette, Physics Reports 478, 1 (2009)
2009
-
[45]
J. P. Garrahan, R. L. Jack, V. Lecomte, E. Pitard, K. van Duijvendijk, and F. van Wijland, Journal of Physics A: Mathematical and Theoretical 42, 075007 (2009)
2009
-
[46]
Den Hollander, Large Deviations, Vol
F. Den Hollander, Large Deviations, Vol. 14 (American Mathematical Soc., 2008)
2008
-
[47]
R. S. Ellis, Entropy, large deviations, and statistical me- chanics (Springer, 2007)
2007
-
[48]
Giardina, J
C. Giardina, J. Kurchan, and L. Peliti, Physical Review Letters 96, 120603 (2006)
2006
-
[49]
U. Ray, G. K.-L. Chan, and D. T. Limmer, Physical Review Letters 120, 210602 (2018)
2018
-
[50]
M. C. Ba˜ nuls and J. P. Garrahan, arXiv preprint arXiv:1903.01570 (2019)
2019 arXiv
-
[51]
Jacobson and S
D. Jacobson and S. Whitelam, Phys. Rev. E 100, 052139 (2019)
2019
-
[52]
D. T. Gillespie, The Journal of Physical Chemistry 81, 2340 (1977)
1977
-
[53]
Seifert, Physical Review Letters 95, 040602 (2005)
U. Seifert, Physical Review Letters 95, 040602 (2005)
2005
-
[54]
Speck, A
T. Speck, A. Engel, and U. Seifert, Journal of Statis- tical Mechanics: Theory and Experiment 2012, P12001 (2012)
2012
-
[55]
Lecomte, A
V. Lecomte, A. Imparato, and F. v. Wijland, Progress of Theoretical Physics Supplement 184, 276 (2010)
2010
-
[56]
Fodor, M
´E. Fodor, M. Guo, N. Gov, P. Visco, D. Weitz, and F. van Wijland, EPL (EuroPhysics Letters) 110, 48005 (2015)
2015
-
[57]
Bucklew, Introduction to rare event simulation (Springer Science & Business Media, 2013)
J. Bucklew, Introduction to rare event simulation (Springer Science & Business Media, 2013)
2013
-
[58]
P. W. Glynn and D. L. Iglehart, Management Science 35, 1367 (1989)
1989
-
[59]
J. S. Sadowsky and J. A. Bucklew, IEEE transactions on Information Theory 36, 579 (1990)
1990
-
[60]
J. A. Bucklew, P. Ney, and J. S. Sadowsky, Journal of Applied Probability 27, 44 (1990)
1990
-
[61]
J. A. Bucklew, Large deviation techniques in decision, simulation, and estimation (Wiley New York, 1990)
1990
-
[62]
Asmussen and P
S. Asmussen and P. W. Glynn, Stochastic simulation: al- gorithms and analysis , Vol. 57 (Springer Science & Busi- ness Media, 2007)
2007
-
[63]
Juneja and P
S. Juneja and P. Shahabuddin, Handbooks in operations research and management science 13, 291 (2006)
2006
-
[64]
Li and F
Y. Li and F. Cao, European Journal of Operational Re- search 224, 333 (2013)
2013
-
[65]
S. S. Singh, V. B. Tadi´ c, and A. Doucet, European Jour- nal of Operational Research 178, 808 (2007)
2007
-
[66]
J. H. Holland, Scientific american 267, 66 (1992)
1992
-
[67]
D. B. Fogel and L. C. Stayton, BioSystems 32, 171 (1994)
1994
-
[68]
Lehman, J
J. Lehman, J. Chen, J. Clune, and K. O. Stanley, in Proceedings of the Genetic and Evolutionary Computa- tion Conference (2018) pp. 450–457
2018
-
[69]
Salimans, J
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever, arXiv preprint arXiv:1703.03864 (2017)
2017 arXiv
- [70]
-
[71]
Lehman, J
J. Lehman, J. Chen, J. Clune, and K. O. Stanley, in Proceedings of the Genetic and Evolutionary Computa- tion Conference (2018) pp. 117–124
2018
-
[72]
Conti, V
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. Stan- ley, and J. Clune, in Advances in neural information processing systems (2018) pp. 5027–5038
2018
-
[73]
T. R. Gingrich, J. M. Horowitz, N. Perunov, and J. L. England, Physical Review Letters 116, 120601 (2016)
2016
-
[74]
J. P. Garrahan, Physical Review E 95, 032134 (2017). 12
2017
-
[75]
Pietzonka, A
P. Pietzonka, A. C. Barato, and U. Seifert, Physical Review E 93, 052145 (2016)
2016
-
[76]
The model’s rates are W12 = 3, W13 = 10, W14 = 9, W21 = 10, W23 = 1, W24 = 2, W31 = 6, W32 = 4, W34 = 1, W41 = 7, W42 = 9, and W43 = 5
-
[77]
Fredrickson and H
G. Fredrickson and H. C. Andersen, Physical Review Let- ters 53, 1244 (1984)
1984
-
[78]
J. P. Garrahan and D. Chandler, Physical Review Letters 89, 035704 (2002)
2002
-
[79]
J. P. Garrahan, R. L. Jack, V. Lecomte, E. Pitard, K. van Duijvendijk, and F. van Wijland, Physical Review Let- ters 98, 195702 (2007)
2007
-
[80]
[51] is more efficient; to calculate the bound itself the choice (w0 or λ) makes little difference
In order to calculate the correction to the bound (3), the parameterization using the variable called λ in Ref. [51] is more efficient; to calculate the bound itself the choice (w0 or λ) makes little difference
-
[81]
Krizhevsky, I
A. Krizhevsky, I. Sutskever, and G. E. Hinton, in Ad- vances in neural information processing systems (2012) pp. 1097–1105
2012
-
[82]
LeCun, L
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, et al., Pro- ceedings of the IEEE 86, 2278 (1998)
1998
-
[83]
P. G. Bolhuis, D. Chandler, C. Dellago, and P. L. Geissler, Annual Review of Physical Chemistry 53, 291 (2002)
2002
-
[84]
R. L. Jack and P. Sollich, The European Physical Journal Special Topics 224, 2351 (2015)
2015
- [85]
-
[86]
Binder, in Monte Carlo Methods in Statistical Physics (Springer, 1986) pp
K. Binder, in Monte Carlo Methods in Statistical Physics (Springer, 1986) pp. 1–45
1986
-
[87]
Maes and K
C. Maes and K. Netoˇ cn` y, EPL (EuroPhysics Letters)82, 30003 (2008)
2008
-
[88]
Bertini, A
L. Bertini, A. Faggionato, D. Gabrielli, et al., in Annales de l’Institut Henri Poincar´ e, Probabilit´ es et Statistiques, Vol. 51 (Institut Henri Poincar´ e, 2015) pp. 867–900
2015
-
[89]
W. K. Hastings, Monte Carlo sampling methods using Markov chains and their applications (Oxford University Press, 1970)
1970
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.