REVIEW 3 major objections 5 minor 2 cited by
Liquid and solid layers in a thermal deep learning machine
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A physics-style training procedure on MNIST shows that an over-parametrized multilayer perceptron develops a solid-liquid-solid layer structure, with layers near the input and output behaving as solids and central layers as liquids.
desk verdict Qualitative solid-liquid-solid structure in trained MLPs is convincingly demonstrated on MNIST, but the quantitative ξ_s ~ ln α phase boundary is not yet established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the thermal deep learning machine, a Hamiltonian model whose energy is the sum over layers of the squared residual between each layer's internal state $S_l$ and the tanh-activated propagated state from the previous layer, with the training data imposed as quenched boundary conditions $S_0$ and $S_L$ and both weights $J_l$ and internal states $S_l$ free to move. Training is performed by GPU-accelerated Langevin molecular dynamics at temperature $T = 10^{-5}$, and the liquid/solid classification is carried by layer-dependent replica correlation functions $q_l(t,t_w)$ and $Q_l(t,t_w)$, which measure the overlap between two independently evolved replicas of the same layer. The key working criterion is that a layer is 'liquid' if the spin replica overlap at the longest simulated time falls below $1/e$, and 'solid' otherwise; this threshold, combined with the storage ratio $\alpha = M/N$, produces the numerical phase diagram of Fig. 1b and the scaling $\xi_s \sim \ln \alpha$.
What would settle it
Run the same MD training of the TDLM for a much longer time (or with a range of thresholds around $1/e$) and recompute the layer-by-layer replica overlap $q^*_l$: if layers currently labeled solid fall below the threshold once the near-boundary layers thermalize, the solid-liquid-solid classification and the $\xi_s \sim \ln \alpha$ scaling would have to be revised. A complementary test is to repeat the experiment on a different real-world dataset and check whether the same phase boundary appears at the same values of $\ln \alpha$.
Extended reading notes
Core claim
Training an over-parametrized fully connected MLP on MNIST with a thermal deep learning machine, the authors find numerical evidence for the theoretically predicted solid-liquid-solid structure: layers adjacent to the input and output behave as solids (slow relaxation, large replica overlaps, hierarchical organization of the design space), while central layers behave as liquids (fast relaxation, near-zero replica overlaps, structureless design space). Using the layer-dependent replica overlap $q^*_l$ measured at $t^* = 5\times10^5$, $t_w = 5\times10^4$ and classified by the threshold $q^*_l = 1/e$, the paper establishes the coexistence of liquid and solid layers and maps the boundary as a function of the storage ratio $\alpha = M/N$. The penetration depth of the solid layers from each boundary grows approximately as $\xi_s \sim \ln \alpha$, consistent with the replica-theory prediction. The paper also documents glassy aging dynamics during training, with logarithmic energy decay and test accuracy growth in the intermediate regime, and shows that beyond a certain depth, adding more layers only extends the liquid region.
Load-bearing premise
The layer classification rests on one threshold and one finite simulation time: a layer is called liquid when $q^*_l < 1/e$ at $t^* = 5\times10^5$ with waiting time $t_w = 5\times10^4$, and the near-boundary solid layers are still out of equilibrium at that moment, so the boundary could shift if longer runs or a different threshold were used.
Editorial extensions
If this is right
- The solid-liquid-solid structure, previously obtained only for random data in replica theory, appears in a realistic benchmark task, so the prediction is not an artifact of structureless inputs.
- Solid penetration depth grows as $\ln(M/N)$, so over-parametrization (small $\alpha$) keeps most layers liquid, while strong over-constraining pushes solid order deep into the network.
- The balance between representation and generalization can be read off the phase diagram: liquid central layers provide flexibility, solid near-boundary layers provide the correlations that yield test accuracy.
- For networks deeper than about $L = 10$ at $\ln \alpha = 5.3$, extra hidden layers are added only to the liquid region; the solid depth stays constant.
- The observed aging regime with logarithmic energy and accuracy growth means that training dynamics themselves carry glassy signatures, not just the static phase structure.
Reading between the lines
- Because the TDLM reduces to a standard MLP as temperature goes to zero, the same solid-liquid-solid layering may appear in gradient-descent-trained networks, but SGD and MD explore the loss landscape differently, so this is a testable conjecture rather than a consequence of the paper.
- The hierarchical overlap structure seen in solid layers resembles the Gardner-transition hierarchy in structural glasses, suggesting that near-boundary layers may store data in a glassy, replica-symmetry-broken manner rather than as a single crystal.
- The supplementary note that MNIST images are correlated implies that the effective number of independent constraints is smaller than $M$; a cleaner version of the phase diagram might use an information-theoretic measure of constraint count instead of raw $M$.
- One practical extension would be to use the layer-resolved overlaps as an early-stopping or layer-pruning diagnostic, since the solid/liquid boundary tells which layers are still underconstrained.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces a 'thermal deep learning machine' (TDLM), a fully connected multilayer perceptron with a Hamiltonian whose dynamical variables are the hidden activations and weights, trained by Langevin molecular dynamics at low temperature on binarized MNIST. The central claim is that the trained network develops a solid-liquid-solid layer structure: layers adjacent to the input and output are solid-like (slow relaxation, high replica overlaps, hierarchical overlap structure), while central layers are liquid-like (fast relaxation, near-zero overlaps). This is summarized in a numerical phase diagram (Fig. 1b) plotted as a function of storage ratio α=M/N and layer index l, with a liquid/solid boundary defined by the replica overlap threshold q*_l < 1/e at a single finite time pair. The paper further claims that the solid penetration depth ξ_s grows as ξ_s ~ ln α, consistent with the theoretical prediction Eq. (1). Global dynamics show aging, and the trained weights generalize to ~98% test accuracy.
Significance. If the central claim is correct, the paper provides the first numerical bridge between the replica mean-field theory of MLPs [14,15] and learning on a real benchmark dataset, and it proposes a physically interpretable training scheme with a well-defined Hamiltonian. The qualitative solid-liquid-solid picture is supported by several complementary and mutually consistent measurements: layer-resolved autocorrelation and relaxation times (Fig. 3a,b), replica overlaps q_l and Q_l (Fig. 3c,d, SI S3), overlap distributions and dendrograms (Fig. 4), and depth dependence (SI S4). The paper is also commendable for not fitting any free coefficient to the target scaling: the comparison with Eq. (1) is a direct test. The main weakness is that the quantitative phase boundary and the ξ_s ~ ln α claim rest on a finite-time overlap threshold applied in layers that are explicitly out of equilibrium.
major comments (3)
- [Layer-dependent training dynamics (Fig. 3b, Fig. 1b)] The liquid/solid classification that underlies Fig. 1b is operational and finite-time: q*_l = q_l(t*=5×10^5, tw=5×10^4) with threshold 1/e. The text explicitly states (Fig. 3b caption) that near-boundary solid layers are still out of equilibrium at this time scale, so q*_l measures whether a layer has decorrelated within the observation window rather than whether its equilibrium Gardner-volume state is liquid or solid. In an aging system, any fixed observation time produces a moving front between 'relaxed' and 'unrelaxed' layers; the growth of the inferred solid depth with ln α could therefore be a dynamical artifact. To make the phase boundary and the ξ_s ~ ln α test convincing, the authors should show that the boundary is stable under variation of (t*, tw) and of the threshold, or replace the single-time criterion by a criterion based on a converged quantity, such as the plateau of τ_c^l for central layers and an extrapolated or equilibrium-based estimate for solid layers.
- [A thermal machine for supervised deep learning; SI Table S1] The quantitative test of Eq. (1) is based on only four values of ln α (3.9, 4.6, 5.3, 6.0) for fixed M=2000, with the width N varied from 40 down to 5; the resulting ξ_s is an integer layer count with no reported error bars and no fit shown. Varying N changes the system size together with α, and the SI (Sec. S3) itself notes that changing N introduces noticeable finite-size effects. With this number of points, the available resolution, and the finite-size confound, the data are at best suggestive of ξ_s ~ ln α rather than a confirmation; the authors should provide a regression with error bars, test at least one independent scaling path (e.g., varying M at fixed larger N), and discuss finite-size corrections.
- [Discussion and phase diagram (Fig. 1b)] The threshold q*_l < 1/e and the evaluation times (t*, tw) are presented as practical definitions, but no justification or sensitivity analysis is given for either choice. Because the phase boundary in Fig. 1b is a central quantitative result, the paper should demonstrate that the boundary location is robust to reasonable variations of these choices, or explicitly characterize how the boundary shifts. Without such an analysis, the comparison with Eq. (1) in the main text and in SI Secs. S3 and S4 remains a comparison to a particular dynamical observable rather than to the equilibrium theory.
minor comments (5)
- [Fig. 1b caption] The caption refers to the 'glass' boundary while the figure color bar and the text use 'solid'; please unify the terminology.
- [Page 3, after Fig. 2] 'standard MLMs' should read 'standard MLPs'.
- [Appendix B, Eq. (2)] Eq. (2) and the surrounding appendix contain typesetting artifacts in the manuscript source (e.g., '/radicaltp'); please ensure the equation is rendered correctly.
- [Appendix E and Fig. 1b] The number of samples Ns and replicas Nr used for each curve in Fig. 1b and Fig. 3 is not stated in the main text; Appendix E gives ranges, but the reader cannot tell which simulation parameters produced the phase diagram. A table analogous to Table S1 for all figures would help.
- [Fig. 2h] The label 'τeq' appears in Fig. 2h before it is defined; please define it explicitly when it first appears.
Circularity Check
No circular reduction found: the numerical experiment is an independent test of a prior replica-theory prediction, and the operational liquid/solid threshold, while arbitrary, is not fitted to the target relation.
full rationale
The paper's central theoretical target, Eq. (1) (ξs ∼ ln α), is imported from Refs. [14,15], which are by coauthor H. Yoshino, but those are prior analytical replica calculations for random inputs. The present paper performs MD simulations on MNIST and compares the observed layer-resolved behavior with that published prediction, so the self-citation states a falsifiable hypothesis rather than replacing the numerical evidence. The liquid/solid assignment is made operationally via q*_l < 1/e at t* = 5×10^5, tw = 5×10^4 (Sec. 'Layer-dependent training dynamics' and Fig. 1b). This is an arbitrary fixed-time threshold, and the paper explicitly states that near-boundary solid layers are still out of equilibrium at this time; consequently, the extracted ξs can in principle depend on the chosen observation window and threshold. That is a robustness limitation on the quantitative test of Eq. (1), but it is not circular: no parameter is fitted to reproduce Eq. (1), and the qualitative solid-liquid-solid picture is independently supported by layer-resolved autocorrelations, relaxation times, overlap distributions, and dendrogram hierarchies. The only other self-reference (Ref. [29], Li and Jin, for logarithmic aging in spin glasses) is illustrative and not load-bearing. No step reduces Eq. (1) or the phase diagram to the input data by construction, and no uniqueness claim from the authors' prior work is used to forbid alternatives. I therefore score 0 despite the presence of self-citations, because none of them performs circular work.
Assumptions & free parameters
free parameters (3)
- liquid/solid threshold q*_l < 1/e =
1/e ≈ 0.368
- overlap evaluation times t* and tw =
t* = 5e5, tw = 5e4
- simulation temperature T =
1e-5
assumptions (4)
- domain assumption The boundary-value Hamiltonian Eq. (2)-(3) with free internal representations {Sl} and weights {Jl} is an appropriate model of supervised learning, reducing to an MLP in the zero-temperature limit.
- domain assumption The theoretical prediction xi_s ~ ln alpha from Refs. [14,15], derived for random structureless inputs in the dense limit, remains valid for structured MNIST data at small widths N <= 40.
- ad hoc to paper Finite-time replica overlap q*_l at t* = 5e5, tw = 5e4 is a valid order parameter for liquid/solid phases, with threshold 1/e.
- domain assumption Weights {Jl} sampled from the thermal ensemble can be evaluated in a sign-activation feed-forward DNN to measure generalization.
invented entities (1)
-
Thermal deep learning machine (TDLM)
Cite this review
Pith. "Pith review of Liquid and solid layers in a thermal deep learning machine." pith.science (2026). https://pith.science/paper/DUW6GOKN
@misc{pith2026250606789,
author = {Pith},
title = {Pith review of: Liquid and solid layers in a thermal deep learning machine},
year = {2026},
howpublished = {\url{https://pith.science/paper/DUW6GOKN}},
note = {Machine review of arXiv:2506.06789}
}
read the original abstract
Based on deep neural networks (DNNs), deep learning has been successfully applied to many problems, but its mechanism is still not well understood -- especially the reason why over-parametrized DNNs can generalize. A recent statistical mechanics theory on supervised learning by a prototypical multi-layer perceptron (MLP) on some artificial learning scenarios predicts that adjustable parameters of over-parametrized MLPs become strongly constrained by the training data close to the input/output boundaries, while the parameters in the center remain largely free, giving rise to a solid-liquid-solid structure. Here we establish this picture, through numerical experiments on benchmark real-world data using a thermal deep learning machine that explores the phase space of the synaptic weights and neurons. The supervised training is implemented by a GPU-accelerated molecular dynamics algorithm, which operates at very low temperatures, and the trained machine exhibits good generalization ability in the test. Global and layer-specific dynamics, with complex non-equilibrium aging behavior, are characterized by time-dependent auto-correlation and replica-correlation functions. Our analyses reveal that the design space of the parameters in the liquid and solid layers are respectively structureless and hierarchical. Our main results are summarized by a data storage ratio -- network depth phase diagram with liquid and solid phases. The proposed thermal machine, which is a physical model with a well-defined Hamiltonian, that reduces to MLP in the zero-temperature limit, can serve as a starting point for physically interpretable deep learning.
Figures
Forward citations
Cited by 2 Pith papers
-
Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation
A replica/HCIZ theory predicts the Bayes-optimal generalization error of proportional-width MLPs near interpolation and discovers layer-wise specialization transitions that make deeper targets harder to learn.
-
Relaxation of quenched structural glasses: descent in a stiffening caging potential over inflection-point 'speed bumps'
Power-law relaxation in gradient-descent glasses is traced to a stiffening caging potential rather than to saddle-point or marginal-stability physics.
Reference graph
Works this paper leans on
-
[1]
8 3.9 4.6 5.3 6.0 solid liquid q∗ l (b) l ln α solid liquid (a) FIG. 1. Liquid-solid phase diagram of deep learn- ing. (a) Schematic phase diagram for the state of net- work parameters in the l-th layer with a given capacity ra- tio α = M/N , drawn according to the theoretical prediction Eq. (1) in N → ∞ limit [14, 15]. (b) Numerical phase dia- gram obtai...
work page 2000
-
[2]
7 0 101 102 103 104 10510− 2 10− 1 100 τth τ0 ln α t E(t)/E (0)
- [3]
-
[4]
5 1 tw t c(t, t w) 102 103 104 101 102 103 104 105
-
[5]
7 0 101 102 103 104 105
-
[6]
8 1 ln α A∗ fixed M = 2000 fixed N = 10 100 101 102 103 104 105 0
work page 2000
-
[7]
5 1 τth tw c∗ C ∗ q∗ Q∗ FIG. 2. Global training dynamics and test accuracy . Time evolution of the (a,c) reduced potential energy E(t)/E (0) and (b,d) test accuracy A(t), for different ln α = ln(M/N ). In (a,b), M = 2000 is fixed and in (c,d) N = 10 is fixed (see Table S1). (e,f) E∗ = E(t∗ ) and A∗ = A(t∗ ), where t∗ = 5 × 105 is the maximum simulation time....
work page 2000
-
[8]
5 1 tw l = 5 l = 9 t cl(t, t w) 103 104 5 × 104 100 101 102 103 104 105 102 103 104 105 106 τth tw τc l 101 102 103 104 1050
Show all 57 references
-
[9]
5 1 l t ql(t, t w = 5 × 104) 1 2 3 4 5 6 7 8 9 1 2 3 4 5 6 7 8 90
-
[10]
8 1 ξl l q∗ l training test FIG. 3. Layer-dependent training dynamics. (a) Spin auto-correlation functions cl(t, t w) in the l = 5 (open symbols) and l = 9 (filled symbols) layers, for tw = 10 3 (black), 10 4 (red), 5 × 104 (blue). (b) Spin correlation time τ c l as a function ...
2000
-
[11]
We set M = 2000 and ln α = 4
and solid ( l = 9) layers. We set M = 2000 and ln α = 4. 6. (a,d) Probability distribution P (Qab l ), for ( Ns, N r) = (20 , 8) (red) and (38 , 16) (black). (b,e) The 36 × 36 overlap matrix of Qab l obtained with ( Ns, N r) = (6 , 6), where the color bar represents the value ...
2000
-
[12]
Monasson and R
R. Monasson and R. Zecchina, Weight space structure and internal representations: a direct approach to learn- ing and generalization in multilayer neural networks, Phys. Rev. Lett. 75, 2432 (1995)
1995
-
[13]
J. J. Hopfield, Neural networks and physical systems with emergent collective computational abilities, Proc. Natl. Aca. Sci. U. S. A. 79, 2554 (1982)
1982
-
[14]
Engel and C
A. Engel and C. Van den Broeck, Statistical mechanics of learning (Cambridge University Press, 2001)
2001
-
[15]
Carleo, I
G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborov´ a, Machine learning and the physical sciences, Rev. Mod. Phys. 91, 045002 (2019)
2019
-
[16]
Nishimori, Statistical physics of spin glasses and in- formation processing: An introduction (Clarendon press, Oxford, 2001)
H. Nishimori, Statistical physics of spin glasses and in- formation processing: An introduction (Clarendon press, Oxford, 2001)
2001
-
[17]
Gardner, Maximum storage capacity in neural net- works, Europhys
E. Gardner, Maximum storage capacity in neural net- works, Europhys. Lett. 4, 481 (1987)
1987
-
[18]
Gardner, The space of interactions in neural network models, J
E. Gardner, The space of interactions in neural network models, J. Phys. A: Math. Gen 22, 257 (1988)
1988
-
[19]
Baldassi, C
C. Baldassi, C. Borgs, J. T. Chayes, A. Ingrosso, C. Lu- cibello, L. Saglietti, and R. Zecchina, Unreasonable ef- fectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes, Proceedings of the National Academy of Sciences ...
2016
-
[20]
Choromanska, M
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun, The loss surfaces of multilayer networks, in Artificial intelligence and statistics (PMLR, 2015) pp. 192–204
2015
-
[21]
A. J. Ballard, R. Das, S. Martiniani, D. Mehta, L. Sagun, J. D. Stevenson, and D. J. Wales, Energy landscapes for machine learning, Physical Chemistry Chemical Physics 19, 12585 (2017)
2017
-
[22]
Baity-Jesi, L
M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G. B. Arous, C. Cammarota, Y. LeCun, M. Wyart, and G. Biroli, Comparing dynamics: Deep neural networks versus glassy systems, Proceedings of the 35th Interna- tional Conference on Machine Learning. PMLR 80, 314 (2018)
2018
-
[23]
Franz, S
S. Franz, S. Hwang, and P. Urbani, Jamming in multi- layer supervised learning models, Physical review letters 123, 160602 (2019)
2019
-
[24]
Geiger, S
M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity- Jesi, G. Biroli, and M. Wyart, Jamming transition as a paradigm to understand the loss landscape of deep neural networks, Physical Review E 100, 012115 (2019)
2019
-
[25]
Yoshino, From complex to simple: hierarchical free- energy landscape renormalized in deep neural networks, SciPost Phys
H. Yoshino, From complex to simple: hierarchical free- energy landscape renormalized in deep neural networks, SciPost Phys. Core 2, 005 (2020)
2020
-
[26]
Yoshino, Spatially heterogeneous learning by a deep student machine, Physical Review Research 5, 033068 (2023)
H. Yoshino, Spatially heterogeneous learning by a deep student machine, Physical Review Research 5, 033068 (2023)
2023
-
[27]
Huang, R
Z.-Y. Huang, R. Zhou, M. Huang, and H.-J. Zhou, Energy-information trade-off induces continuous and dis- continuous phase transitions in lateral predictive coding, Science China Physics, Mechanics & Astronomy 67, 1 (2024)
2024
-
[28]
M. K. Winter and L. Janssen, How glassy are neural net- works?, arXiv preprint arXiv:2405.13098 (2024)
2024 arXiv
-
[29]
Nelson, T
D. Nelson, T. Piran, and S. Weinberg, Statistical mechan- ics of membranes and surfaces (World Scientific, 2004)
2004
-
[30]
Krzakala and L
F. Krzakala and L. Zdeborov´ a, On melting dynamics and the glass transition. i. glassy aspects of melting dynamics, The Journal of chemical physics 134, 034512 (2011)
2011
-
[31]
Krzakala and L
F. Krzakala and L. Zdeborov´ a, On melting dynamics and the glass transition. ii. glassy dynamics as a melting pro- cess, The Journal of chemical physics 134, 034513 (2011)
2011
-
[32]
Biroli, J.-P
G. Biroli, J.-P. Bouchaud, A. Cavagna, T. S. Grigera, and P. Verrocchio, Thermodynamic signature of growing amorphous order in glass-forming liquids, Nature Physics 4, 771 (2008)
2008
-
[33]
W. Kob, S. Rold´ an-Vargas, and L. Berthier, Non- monotonic temperature evolution of dynamic correlations in glass-forming liquids, Nature Physics 8, 164 (2012)
2012
-
[34]
G. M. Hocky, T. E. Markland, and D. R. Reichman, Growing point-to-set length scale correlates with grow- ing relaxation times in model supercooled liquids, Phys- ical review letters 108, 225506 (2012)
2012
-
[35]
Ikeda and A
H. Ikeda and A. Ikeda, One-dimensional kac model of dense amorphous hard spheres, EPL (Europhysics Let- ters) 111, 40007 (2015)
2015
-
[36]
Singh, M
S. Singh, M. D. Ediger, and J. J. De Pablo, Ultrastable glasses from in silico vapour deposition, Nature materials 6 12, 139 (2013)
2013
-
[37]
Berthier, P
L. Berthier, P. Charbonneau, E. Flenner, and F. Zam- poni, Origin of ultrastability in vapor-deposited glasses, Physical review letters 119, 188002 (2017)
2017
-
[38]
LeCun, C
Y. LeCun, C. Cortes, and C. Burges, MNIST handwrit- ten digit database
-
[39]
Qiao, The MNIST database of handwritten digits, Re- trieved 18 August 2013 (2007)
Y. Qiao, The MNIST database of handwritten digits, Re- trieved 18 August 2013 (2007)
2007
-
[40]
Li and Y
B. Li and Y. Jin, The simplest spin glass revisited: finite - size effects of the energy landscape can modify aging dynamics in the thermodynamic limit, arXiv preprint arXiv:2501.00338 (2024)
2024
-
[41]
Hastie, R
T. Hastie, R. Tibshirani, and J. Friedman, The ele- ments of statistical learning , Springer Series in Statistics (Springer, New York, NY, 2009)
2009
-
[42]
Charbonneau, J
P. Charbonneau, J. Kurchan, G. Parisi, P. Urbani, and F. Zamponi, Fractal free energy landscapes in structural glasses, Nature communications 5, 3725 (2014)
2014
-
[43]
Charbonneau, Y
P. Charbonneau, Y. Jin, G. Parisi, C. Rainone, B. Seoane, and F. Zamponi, Numerical detection of the gardner transition in a mean-field glass former, Physical Review E 92, 012316 (2015)
2015
-
[44]
Berthier, P
L. Berthier, P. Charbonneau, Y. Jin, G. Parisi, B. Seoane, and F. Zamponi, Growing timescales and lengthscales characterizing vibrations of amorphous solids, Proceedings of the National Academy of Sciences 113, 8397 (2016)
2016
-
[45]
Urbani, Y
P. Urbani, Y. Jin, and H. Yoshino, The gardner glass, in Spin Glass Theory and Far Beyond: Replica Symmetry Breaking After 40 Years (World Scientific, 2023) pp. 219– 238
2023
-
[46]
Paszke, S
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, Automatic differentiation in pytorch, (2017)
2017
-
[47]
0” and “1
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haber- land, T. Reddy, D. Cournapeau, E. Burovski, P. Peter- son, W. Weckesser, J. Bright, et al., Scipy 1.0: fundamen- tal algorithms for scientific computing in python, Nature methods 17, 261 (2020). 7 End Matter Appendix A: Input d...
2020
-
[48]
5 1 t cl(t, t w) 101 102 103 104 1050
-
[49]
5 1 t ql(t, t w) 101 102 103 104 1050
-
[50]
5 1 ln α t c8(t, t w) 3.9 4.6 5.3 6.0 101 102 103 104 1050
-
[51]
5 1 t q8(t, t w) 101 102 103 104 1050
-
[52]
5 1 l t Cl(t, t w) 1 2 3 4 5 6 7 8 9 10 101 102 103 104 1050
-
[53]
5 1 t Ql(t, t w) 101 102 103 104 1050
-
[54]
5 1 t C8(t, t w) 101 102 103 104 1050
-
[55]
5 1 t Q8(t, t w) FIG. S3. Layer-specific correlation functions. Data are obtained with a fixed waiting time tw = 5 × 104. (a) Spin auto- correlation functions cl(t, t w), (b) weight auto-correlation functions Cl(t, t w), (c) spin replica-correlation functions ql(t, t w), and (d)...
2000
-
[56]
Yoshino, From complex to simple: hierarchical free-energy l andscape renormalized in deep neural networks, SciPost Phys
H. Yoshino, From complex to simple: hierarchical free-energy l andscape renormalized in deep neural networks, SciPost Phys. Core 2, 005 (2020)
2020
-
[57]
Yoshino, Spatially heterogeneous learning by a deep stu dent machine, Physical Review Research 5, 033068 (2023)
H. Yoshino, Spatially heterogeneous learning by a deep stu dent machine, Physical Review Research 5, 033068 (2023)
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.