REVIEW 3 major objections 4 minor 19 references
Typicality of thermal states in isolated quantum systems corresponds to ubiquity of global minima in wide artificial neural networks
T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read This paper claims that quantum thermalization and wide-neural-network training are two views of the same statistical mechanism: when only a few observables matter, almost all microstates look alike.
desk verdict A suggestive, well-informed analogy between canonical typicality and NTK overparameterization, but the load-bearing mapping and spectral scaling are asserted rather than established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Wishart-type matrix. On the quantum side, the reduced density matrix of a Haar-random pure state, expressed in a product basis, is a normalized Wishart matrix whose spectrum follows the Marchenko–Pastur distribution; its deviation from the microcanonical ensemble is controlled by the ratio c = d_A/d_B. On the network side, the Neural Tangent Kernel K = J J^T is viewed as a Wishart matrix of output Jacobians J of dimension (n_L N) × P, whose smallest eigenvalue vanishes when P equals κ n_L N, marking the fitting threshold. The paper uses these two spectral objects to show that the Page point d_A = d_B and the fitting threshold obey the same condition under the dimens
What would settle it
A numerical experiment that measures the smallest eigenvalue of the NTK kernel in finite-width networks and simultaneously computes the entanglement entropy of typical pure states in finite-dimensional Hilbert spaces: if the fitting threshold (where the smallest eigenvalue vanishes) occurs at a parameter count that does not scale like d_A d_B under the proposed mapping, or if the Page peak occurs at a subsystem size where the NTK does not exhibit a double-descent peak, the claimed correspondence fails. More directly, one could test whether the Rényi entanglement entropy of typical pure states
Extended reading notes
Core claim
The paper's central claim is that the typicality of pure thermal states — the fact that for a fixed observable the expectation value in a Haar-random energy-shell state almost always agrees with the microcanonical average, with deviations bounded by a variance-over-dimension inequality — corresponds directly to the ubiquity of global minima in the Neural Tangent Kernel regime. The correspondence is mediated by identifying the total energy-shell dimension d_A d_B with the number of parameters P and the number of linearly independent observables on the subsystem d_A^2 with the number of scalar outputs n_L N. Just as canonical typicality makes a reduced density matrix indistinguishable from the
Load-bearing premise
The load-bearing premise is the identification of the energy-shell dimension d_A d_B with the number of network parameters P and the subsystem observable count d_A^2 with the number of scalar outputs n_L N; this mapping is asserted rather than derived, and the claimed correspondence between the Page point and the fitting threshold depends entirely on it.
Editorial extensions
If this is right
- If the correspondence is correct, thermalization in isolated quantum systems and successful training of wide neural networks share a single mathematical origin: both rely on overparameterization making the few observables of interest insensitive to the microscopic state.
- The threshold at which a small subsystem's reduced state becomes indistinguishable from the thermal ensemble is determined by the same condition as the fitting threshold in NTK, so results about one problem can be translated into predictions about the other.
- The peak of the entanglement entropy at d_A = d_B corresponds to the peak of generalization error at the interpolation threshold; beyond that point both quantities decrease again, implying that 'too much' entanglement or overparameterization restores good behavior.
- The Marchenko–Pastur law provides the quantitative bridge: the spectral edge of the reduced density matrix and of the NTK kernel determine both the thermal deviation and the fitting threshold.
- This extends the previously noted black-hole-evaporation/quantum-machine-learning connection to a more general, purely quantum-thermodynamic setting.
Reading between the lines
- If the correspondence is taken seriously, it suggests a research program where quantum thermalization phenomena (e.g., eigenstate thermalization or many-body localization) are used to predict regimes where neural networks will or will not generalize, and vice versa; this is the author's stated goal but not yet demonstrated.
- The mapping d_A^2 ↔ n_L N may imply that the relevant 'observable dimension' in a neural network is quadratic in the output width, so width must grow like the square root of the number of parameters to match subsystem behavior; a testable consequence is that the double descent peak should shift with n_L rather than with the total parameter count alone.
- One could test the correspondence by measuring the eigenvalue distribution of the empirical NTK kernel in a finite-width network and comparing it with the Marchenko–Pastur prediction of the reduced density matrix: the same spectral-edge scaling should describe both the fitting threshold and the thermalization threshold.
- The correspondence remains qualitative at the level of 'structural similarity,' so a natural extension is to find a single random-matrix ensemble that interpolates between the quantum reduced-density-matrix setting and the NTK kernel setting, which would promote the analogy to a derivation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that two well-known high-dimensional phenomena—the typicality of thermal pure states in isolated quantum systems and the ubiquity of low-training-error global minima in sufficiently wide neural networks—share a common structural mechanism. The proposed mechanism is the combination of a high-dimensional ambient space and the restriction of attention to a small number of observables or outputs, with a Wishart-type random matrix playing a central role. The first half maps the energy-shell dimension and the subsystem dimension to the number of network parameters and training observables; the second half identifies the Page point d_A = d_B with the NTK fitting threshold P = κ n_L N, and interprets the entropy curve as a counterpart of double descent. The paper is explicitly presented not as a detailed mathematical derivation but as a structural correspondence.
Significance. The claimed correspondence is suggestive and, if made rigorous or even sharply formulated as a falsifiable analogy, could connect two active research areas in a useful way. The paper builds on standard ingredients—Reimann's typicality bound, Marchenko–Pastur spectral density, exact Page/Lubkin entanglement expressions, and NTK linearization—and the side-by-side table in Sec. I is helpful. However, the central quantitative bridge is asserted rather than established, and as written the claims outrun the evidence. The paper would be a more valuable contribution if it either supplied a rigorous derivation for the load-bearing spectral statement or transparently reframed the results as a heuristic analogy with explicit conjectures.
major comments (3)
- [Sec. III, paragraph after Eq. (3)] The smallest-eigenvalue scaling lambda_- ≈ sigma(√P - κ√(n_L N)) is asserted for K = JJ^T even though the text acknowledges that the entries of K are not i.i.d. The cited reference [19] is a general non-asymptotic random matrix survey; it does not establish Marchenko–Pastur edge universality for deep NTK Jacobians whose columns are strongly correlated through shared weights and activations. This spectral edge is load-bearing: it yields the fitting threshold P = κ n_L N and the claimed Page-point correspondence. Without a proof or a precisely stated conjecture with supporting numerics, the second main result is not established.
- [Sec. III, mapping d_A d_B ↔ P and d_A^2 ↔ N] The mapping between quantum dimensions and network quantities is asserted, not derived, and it is internally inconsistent when made quantitative. If N corresponds to d_A^2, then the threshold P = κ n_L N becomes d_A d_B = κ n_L d_A^2, i.e. d_B/d_A = κ n_L, not d_A = d_B. If instead n_L N corresponds to d_A^2, then N is no longer the number of observables as claimed. Either way, the stated identification does not imply that the Page point coincides with the fitting threshold. This inconsistency undermines the claimed 'direct correspondence' between the Page curve and double descent.
- [Secs. II–III, overall logical status] The central claim that typicality 'corresponds' to ubiquity of global minima is never given a precise formal meaning. Statements such as 'we show' and 'we demonstrate' are used, but the argument reduces to a list of structural similarities (high dimensionality, few observables, Wishart matrices). With no defined mapping between quantum observables and neural outputs, and no shared mathematical statement that is proved, the correspondence is difficult to falsify. The authors should either formulate a concrete mathematical assertion (e.g., a probabilistic bound in one setting that transfers to the other) or explicitly label the paper as proposing an analogy and conjecture.
minor comments (4)
- [Abstract and Sec. II first paragraph] The phrase 'pure thermal thermal states' contains a duplicated word. Please correct.
- [Eq. (2)] The relation between eigenvalues of the reduced density matrix and the Marchenko–Pastur distribution needs normalization. As written, the support (1±√c)^2 has unit mean, whereas ρ_A has trace one; the scaling by 1/d_A should be stated explicitly.
- [Sec. III, definition of N] The text uses N for the number of input data points and also treats 'number of observables' as interchangeable with N. Since each output is n_L-dimensional, the total number of scalar observables is n_L N. This ambiguity contributes to the mapping inconsistency noted above and should be clarified.
- [References] Reference [19] is a general introduction to non-asymptotic random matrix theory. Please cite the specific theorem or section that is intended to support the edge-scaling claim, or replace it with a result on structured Wishart matrices.
Circularity Check
The claimed Page-point/fitting-threshold correspondence is imposed by the dimension dictionary rather than derived.
-
self definitional
[Sec. III (Distinguishability and double descent), paragraphs defining the dimension dictionary and the fitting-threshold equivalence]
"First, we point out that the roles of fixed quantities and variables are interchanged in the Page curve and double descent phenomenon. In fact, the total dimension d_A d_B is fixed, while the number of linearly independent observables on subsystem A varies as d_A changes. The total dimension of the energy shell d_A d_B corresponds to the number of parameters in ANN. Also, the dimension of the subsystem d_A determines the number of linearly independent observables on subsystem A, which is d^2_A. ... Note that the entanglement entropy S attains its maximum at d_A = d_B, which corresponds to the"
The paper presents as a result that the Page maximum d_A=d_B corresponds to the ANN fitting threshold P=κ n_L N. But under the dictionary just asserted, P is identified with d_A d_B and n_L N is identified with d_A^2. The fitting-threshold condition P=κ n_L N is then d_A d_B = κ d_A^2, i.e. d_B=κ d_A; with κ set to unity it is exactly the Page-point condition. Thus the claimed equivalence of the two thresholds is not a derived correspondence but is chosen by the dimension identifications. The 'second main result' is therefore a renaming of Page's maximum under a coordinate dictionary rather than an independent prediction. Since the same dictionary is the only quantitative bridge between the two subjects, the argument is circular at this load-bearing step.
full rationale
The paper does not rely on self-citation: its two factual pillars—quantum typicality (Goldstein et al., Popescu et al., Reimann, Sugita) and the NTK global-minimum theorem (Jacot et al.)—are external and independently established. The first main result (typicality ↔ ubiquity of global minima) is a qualitative structural analogy supported by those independent results, so it is not circular. The circularity is concentrated in Sec. III. The paper asserts the dictionary d_A d_B ↔ P and d_A^2 ↔ n_L N, then announces that the Page maximum d_A=d_B 'corresponds to the fitting threshold' P=κ n_L N. Under that dictionary the two conditions are the same condition up to an undetermined constant. Thus the second main result—the correspondence between state distinguishability/Page curve and double descent—is a renaming/definitional identification rather than a derived prediction. The paper's additional claim that the smallest eigenvalue of the NTK Gram matrix K=JJ^T follows Marchenko-Pastur edge scaling even though the entries are not independent is an unverified technical assumption (a correctness risk), not a circularity; it is cited to an external source and, if wrong, only weakens the analogy. Overall, the central quantitative bridge is constructed by definition, so partial circularity (score 6) is appropriate.
Assumptions & free parameters
free parameters (1)
- σ and κ in spectral edge scaling =
not determined
assumptions (4)
- domain assumption Typicality inequality (Reimann) and Haar-uniform sampling on the energy shell.
- standard math Marchenko–Pastur law for the eigenvalues of a Wishart matrix.
- domain assumption NTK regime: infinite-width network, linearized training, Lipschitz continuity.
- ad hoc to paper Mapping between quantum dimensions and neural network quantities: d_A d_B ↔ P and d_A^2 ↔ N.
Cite this review
Pith. "Pith review of Typicality of thermal states in isolated quantum systems corresponds to ubiquity of global minima in wide artificial neural networks." pith.science (2026). https://pith.science/paper/QYI6YSPB
@misc{pith2026251109526,
author = {Pith},
title = {Pith review of: Typicality of thermal states in isolated quantum systems corresponds to ubiquity of global minima in wide artificial neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/QYI6YSPB}},
note = {Machine review of arXiv:2511.09526}
}
read the original abstract
The Neural Tangent Kernel theory theoretically guarantees the existence of global minima of the cost functional in the neighborhood of an arbitrary random initialized parameters in wide artificial neural networks. In this paper, we show that the ubiquity of the global minima directly corresponds to the typicality of pure thermal states in isolated quantum systems by identifying a common underlying mechanism characterized by the restriction to a few observables and the role of a Wishart-type matrix. Moreover, we demonstrate that the increase in distinguishability of the reduced density matrices of typical pure states with subsystem size corresponds to the double descent phenomenon observed by varying the width of layers in finite-width artificial neural networks. Thereby, the threshold for the reduced state become thermal is determined by essentially the same condition as the fitting threshold. In this manner, we reveal a structural correspondence between thermalization in isolated quantum systems and wide neural network.
Reference graph
Works this paper leans on
- [19]
-
[1]
Jacot, F Gabriel, and Cl´ ement Hongler, Advances in neural information processing systems, 8571-8580 (2018)
A. Jacot, F Gabriel, and Cl´ ement Hongler, Advances in neural information processing systems, 8571-8580 (2018)
2018
-
[2]
J. W. Lee and Z. Y. Kim, arxiv: 2506.09678
-
[3]
Goldstein, J
S. Goldstein, J. L. Lebowitz, R. Tumulka, and N. Zangh ` ı, Canonical Typicality, Phys. Rev. Lett.96, 050403 (2006)
2006
-
[4]
Popescu, A
S. Popescu, A. J. Short, and A. Winter, Nat. Phys.2, 754 (2006)
2006
-
[5]
A. Sugita, Nonlinear Phenom. Complex Syst.10, 192 (2007); arXiv:cond-mat/0602625
arXiv 2007
-
[6]
Reimann, Phys
P. Reimann, Phys. Rev. Lett.99, 160404 (2007)
2007
-
[7]
R. V. Jensen and R. Shankar, Phys. Rev. Lett.54.1879 (1985)
1985
Show all 19 references
-
[8]
Rigol, V
M. Rigol, V. Dunjko, and M. Olshanii, Nature (London) 452, 854 (2008)
2008
-
[9]
Gring, M
M. Gring, M. Kuhnert, T. Langen, T. Kitagawa, B. Rauer, M. Schreitl, I. Mazets, D. A. Smith, E. Demler, and J. Schmiedmayer, Science337, 1318 (2012)
2012
-
[10]
Trotzky, Y
S. Trotzky, Y. A. Chen, A. Flesch, I. P. McCulloch, U. Schollw¨ ock, J. Eisert, and I. Bloch, Nat. Phys.8, 325 (2012)
2012
-
[11]
von Neumann, Z
J. von Neumann, Z. Phys.57, 30 (1929)
1929
-
[12]
Goldstein, J
S. Goldstein, J. L. Lebowitz, C. Mastrodonato, R. Tu- mulka, and N. Zangh ` ı, Proc. R. Soc. London A466, 3203 (2010)
2010
-
[13]
Zyczkowski,and H
K. Zyczkowski,and H. J. Sommers, J. Phys. A34, 7111 (2001)
2001
-
[14]
V. A. Marˇ cenko, L. A. Pastur, Math. USSR-Sb.1, 457 (1967)
1967
-
[15]
Bai and J
Z. Bai and J. W. Silverstein,Spectral Analysis of Large Dimensional Ran dom Matrices. Springer Series in Statistics, Springer, New York, 2nd edition, (2010)
2010
-
[16]
Cheng, and A
X. Cheng, and A. Singer, Random Matrices: Theory and Applications, 02(04), 1350010 (2013)
2013
-
[17]
Lubkin, J
E. Lubkin, J. Math. Phys.19, 1028 (1978)
1978
-
[18]
D. N. Page, Phys. Rev. Lett.71, 1291 (1993)
1993
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.