REVIEW 4 major objections 4 minor 1 cited by
Black hole/quantum machine learning correspondence
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that the Page time of a black hole corresponds to the interpolation threshold of a quantum linear regression, where the test-error variance diverges and falls on both sides.
desk verdict Clever dictionary, but the central variance formula is unsupported because D is a constant matrix for normalized microstates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the Marchenko-Pastur law, the universal eigenvalue distribution of large random covariance matrices, applied to the reduced density matrix $\rho_r$ of the Hawking radiation and to the data matrix $D$ whose entries are the diagonal overlaps $\langle \psi_i^\beta|\psi_i^\beta\rangle/P$. The ratio $\alpha = \Omega/e^S = P/N$ sets the scale: for $\alpha<1$ the Gram-like matrix is full rank and the least-squares solution is unique, while for $\alpha>1$ zero modes appear and the minimum-norm solution takes over. The variance computation reduces to the Stieltjes transform $S(0) = 1/(1-\alpha)$ of the MP distribution, giving the two diverging-variance formulas, and the projection operator $\Pi = D^\dagger(DD^\dagger)^{-1}D$ characterizes when a test state lies in the recoverable radiation subspace.
What would settle it
Construct a microscopic model of the black hole microstates, for example Haar-random states or a concrete holographic model, and compute the eigenvalue distribution of $D = (1/P)\langle \psi_i^\beta|\psi_i^\beta\rangle$. If the Stieltjes transform at zero does not approach $1/(1-\alpha)$, or if the finite-$N$ variance of the quantum linear regression does not diverge as $\alpha$ approaches 1, the claimed correspondence fails.
Extended reading notes
Core claim
The central claim is that the Page curve of a black hole is governed by the same mathematics as the generalization error of quantum linear regression, with the ratio $\alpha = \Omega/e^S = P/N$ playing the role of the parameter-to-sample ratio. Using the Marchenko-Pastur law for the eigenvalue distribution of the reduced density matrix of the radiation, the authors derive the variance of the regression error as $V = \sigma^2 \alpha/(1-\alpha)$ for $\alpha<1$ and $V = \sigma^2/(\alpha-1)$ for $\alpha>1$, diverging at $\alpha=1$ and decreasing on either side. They interpret this divergence as a quantum phase transition at the Page time, and note that the $\alpha \leftrightarrow 1/\alpha$ inversion symmetry of the Page curve is exactly the symmetry that produces the second descent in overparameterized learning. The conclusion is that the information in a test radiation state is fully recoverable from the radiation subsystem alone only after the Page time.
Load-bearing premise
The derivation assumes that the matrix $D$, built from the diagonal overlaps $\langle \psi_i^\beta|\psi_i^\beta\rangle$, has an eigenvalue spectrum described by the Marchenko-Pastur law; the paper does not prove that $D$ is a random matrix of the required type, nor that the variance formula $\sigma^2 \alpha/(1-\alpha)$ actually holds for the specific $D$ of the radiation state.
Editorial extensions
If this is right
- For $\alpha<1$, the radiation density matrix is full rank and the regression is underparameterized; the prediction variance grows as $\alpha/(1-\alpha)$, reproducing the first descent of the double descent curve.
- For $\alpha>1$, the radiation density matrix becomes rank-deficient with zero modes, the minimum-norm solution applies, and the variance falls as $1/(\alpha-1)$, producing the second descent.
- The Page time is identified with the interpolation threshold $\alpha=1$, where the variance diverges and information recovery undergoes a phase transition.
- A test radiation state can be faithfully reconstructed from the radiation subsystem only after the Page time, because only then does the projection operator $\Pi$ cover the full test-state subspace.
Reading between the lines
- If the correspondence holds, then black hole models with different spectral distributions than Marchenko-Pastur would predict modified double-descent curves, giving a testable fingerprint of the microstate geometry.
- The inversion symmetry $\alpha \leftrightarrow 1/\alpha$ might be a general duality between the interior and exterior descriptions, suggesting the same symmetry underlies both Page's result and the good-generalization regime of overparameterized models.
- One could test the correspondence indirectly by engineering quantum many-body systems whose entanglement dynamics mimics the Page curve, then performing quantum linear regression on their reduced density matrices to look for the predicted divergence in the learned-observable variance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a correspondence between black hole evaporation and the double descent phenomenon in quantum machine learning. It identifies the ratio α = Ω/e^S of radiation Hilbert space dimension to black hole microstate count with the parameter-over-sample ratio P/N of a quantum linear regression, derives a variance formula V = σ²α/(1−α) for α<1 and V = σ²/(α−1) for α>1 using the Marchenko–Pastur law, and concludes that the Page time corresponds to the interpolation threshold, with test information fully recoverable from the radiation only after the Page time. Section II reproduces the Page curve from an MP-type spectral density for the radiation density matrix, and Section III attempts to apply random matrix resolvent techniques to a data matrix D built from diagonal overlaps of black hole microstates.
Significance. If the central derivation were valid, the paper would offer a conceptually appealing bridge between black hole information theory and statistical learning theory, and the identification of the Page time with a divergence in regression variance would be a striking quantitative claim. The paper also correctly recalls standard MP-based derivations of the Page curve and cites relevant prior work on black hole spectral statistics. However, the load-bearing random matrix calculation in Section III is invalid as written, because the matrix D defined in Eq. (14) is constant for normalized microstates and cannot be governed by the Marchenko–Pastur law. The central quantitative result is therefore unsupported, and the correspondence reduces to a parameter identification.
major comments (4)
- [Section III, Eq. (14)–(19)] The definition D_{βi} = ⟨ψ_i^β|ψ_i^β⟩/P in Eq. (14), combined with the normalization of the states in Eq. (1), gives ⟨ψ_i^β|ψ_i^β⟩ = 1, so every entry of D equals 1/P. Thus D is the constant matrix (1/P) 1_{N×P}, D^†D = (N/P^2) J_P is a rank-one all-ones matrix, and Tr((D^†D)^{-1}) is undefined for every α. The MP-based limit V = σ²α/(1−α) in Eq. (19) cannot follow from this D. The MP law in Eq. (10) applies to the full Gram matrix ⟨ψ_i^β|ψ_j^β⟩ appearing in ρ_r, not to the diagonal-overlap matrix D. This invalidates the main quantitative claim that the variance diverges at the Page time α=1.
- [Section III, projector and recovery argument] The projection operator Π = D^†(DD^†)^{-1}D is introduced to discuss full recovery of the test state. For the actual D defined in Eq. (14), both D^†D and DD^† have rank one for all α, so the inverses appearing in Π do not exist in the matrix sense, and Π is not a projector onto a subspace that changes rank at α=1. Consequently, the claim that 'this condition becomes feasible only after the Page time' is not supported by the model as written. The recovery argument would require a different, justified ensemble for D.
- [Section III, Eqs. (16)–(20) and Fig. 1] The derivation computes only the variance term in the regression error, yet the paper repeatedly refers to 'test error' and plots a divergence in Fig. 1 as the double descent phenomenon. In linear regression the total test error also contains a bias term, which is not computed. Even if the variance formula were correct, the statements about the test error exhibiting double descent would remain incomplete without a treatment of the bias.
- [Section III, after Eq. (17)] The assertion that 'the above error approaches σ²αS(0) for D associated with ρ_r in Eq. (2)' is not justified. No argument is supplied showing that D, or any matrix constructed from the diagonal overlaps, has an eigenvalue distribution converging to the Marchenko–Pastur law. The entries of D are deterministic given the normalized microstates, not i.i.d. random variables as required for Eq. (10). If D were instead replaced by a generic random matrix, the connection to the black hole microstate model would be lost. A precise random-matrix ensemble for D is needed, and none is given.
minor comments (4)
- [Section II, Eq. (3)] The expression for f(λ) contains unbalanced parentheses in the square-root factor; the formula should be typeset with the two bracketed terms (λ−(Ω^{1/2}−e^{−S/2})²) and ((Ω^{1/2}+e^{−S/2})²−λ) displayed separately.
- [Section III, paragraph after Eq. (13)] The phrase 'the first term (the variance)' is not tied to a displayed formula; the split into bias and variance terms appears only implicitly in the later residual decomposition. A clearer separation of the two terms would help the reader follow which quantity is being computed.
- [Section III, Eq. (14)] The notation D is used both for the N×P data matrix and, implicitly, for the density matrix ρ_r. The distinction between D and the Gram matrix of overlaps should be made explicit, since the current notation invites confusion between diagonal overlaps and the off-diagonal overlaps that actually enter the MP distribution.
- [Section III, recovery paragraph] The statement that 'the information contained in ρ_t is fully recoverable from the radiation subsystem alone only after the Page time' is a strong claim, but the paragraph does not specify the sense in which recovery is possible (exact operator reconstruction, estimation with finite samples, or information-theoretic distinguishability). This precision is needed because the mathematical condition involving Π is not proved.
Circularity Check
The central variance formula is the classical Marchenko–Pastur double-descent result relabeled with black-hole variables; the assertion that D obeys the MP law is unproven and inconsistent with normalized microstates.
-
renaming known result
[Section III, Eq. (19) and Eq. (17); compare Appendix A, Eq. (A11)]
"In the limit P=αN→∞, the above error approaches σ2αS(0) [9] for D associated with ρr in Eq. 2, where S(z)≡∫ (1/(λ−z)) fMP(λ)dλ is the Stieltjes transform. Using the resolvent method one can obtain S(0)=1/(1−α) and then V=Eϵ[∥W*−Wunder∥2]=σ2α/(1−α), which diverges at the interpolation threshold (α=1)."
This is the load-bearing quantitative result. It imports Eq. (A11), the classical MP variance for an i.i.d. Gaussian design matrix X, and applies it verbatim to the matrix D defined in Eq. (14), after substituting P=Ω and N=e^S. No argument shows that D has the MP spectrum; in fact, for normalized microstates ⟨ψ_i^β|ψ_i^β⟩=1, so every entry of D is 1/P and D^†D is rank-one, not MP. Thus the 'Page-time divergence' is the known MP result relabeled, not a derived consequence of the black-hole model.
-
self definitional
[Section II (parameter identification) and Section III (paragraph containing Eq. (19))]
"By identifying P=Ω and N=eS, one can draw a correspondence between QML and the BH physics. ... Therefore, the test error diverges at the interpolation threshold (i.e., the Page time)."
Because α was defined as Ω/e^S and then equated to P/N, the statement that the Page time (Ω≃e^S) occurs at the interpolation threshold (P=N) is true by definition. Presenting this coincidence as a derived correspondence, and locating the variance divergence at the Page time, relies on the dictionary rather than on a calculation. The only nontrivial content remaining in the divergence is the MP law imported in the previous step.
full rationale
The paper does not rely on self-citation; its references are external and legitimate. The circularity is of a different kind: the central quantitative claim, V=σ²α/(1−α) with divergence at α=1, is the standard Marchenko–Pastur double-descent formula from Appendix A (Eq. A11), restated after renaming P=Ω, N=e^S, and α=P/N. The step 'for D associated with ρr in Eq. 2' is asserted without proof, and for normalized microstates the diagonal-overlap matrix D is constant and rank-one, so the MP spectrum cannot be assumed for it. Consequently, the claimed derivation of a Page-time divergence reduces to an identification of variables plus an imported random-matrix result, rather than a consequence derived from the black-hole model itself. The qualitative rank transition (α<1 full-rank vs α>1 rank-deficient) is independently known from the Page curve, but the double-descent variance curve is not independently derived here. Score 6 reflects partial circularity: one or more predictions reduce by construction to the input MP law under the imposed dictionary.
Assumptions & free parameters
free parameters (1)
- sigma (noise variance) =
1 (set for Fig. 1)
assumptions (5)
- domain assumption The eigenvalue density of the reduced density matrix rho_r follows the Marchenko-Pastur form in Eq. (3).
- ad hoc to paper The predictor matrix D has a Marchenko-Pastur distributed eigenspectrum.
- domain assumption There exist N = e^S degenerate internal states labeled by beta = 1,...,N.
- domain assumption The Page curve formula S_r = S + log(alpha) - alpha/2 for alpha <= 1 and S - 1/(2 alpha) for alpha >= 1 is correct.
- standard math Standard least squares solutions and the Stieltjes transform / resolvent method apply as in classical linear regression.
Cite this review
Pith. "Pith review of Black hole/quantum machine learning correspondence." pith.science (2026). https://pith.science/paper/7WFGMRQJ
@misc{pith2026250609678,
author = {Pith},
title = {Pith review of: Black hole/quantum machine learning correspondence},
year = {2026},
howpublished = {\url{https://pith.science/paper/7WFGMRQJ}},
note = {Machine review of arXiv:2506.09678}
}
read the original abstract
We explore a potential connection between the black hole information paradox and the double descent phenomenon in quantum machine learning. Information retrieval from Hawking radiation can be viewed through the lens of quantum linear regression over black hole microstates, with the Page time corresponding to the interpolation threshold, beyond which test error decreases despite overparameterization. Using the Marchenko-Pastur law, we derive the variance in test error for the quantum linear regression problem and show that the transition across the Page time is associated with a change in the rank structure of subsystems. This observation suggests a conceptual parallel between black hole physics and machine learning that may provide new perspectives for both fields.
Figures
Forward citations
Cited by 1 Pith paper
-
Typicality of thermal states in isolated quantum systems corresponds to ubiquity of global minima in wide artificial neural networks
A structural correspondence is drawn between canonical typicality in isolated quantum systems and the NTK genesis of global minima in wide neural networks, with the Page point mapped to the fitting threshold.
Reference graph
Works this paper leans on
-
[1]
A. Almheiri, T. Hartman, J. Maldacena, E. Shaghoulian, and A. Tajdini, Rev. Mod. Phys.93, 035002 (2021), 2006.06872
arXiv 2021
-
[2]
S. W. Hawking, Phys. Rev. D14, 2460 (1976)
1976
-
[3]
D. N. Page, Phys. Rev. Lett.71, 1291 (1993)
1993
-
[4]
D. N. Page, Phys. Rev. Lett.71, 3743 (1993)
1993
- [5]
-
[6]
M. Loog, T. Viering, A. Mey, J. H. Krijthe, and D. M. J. Tax, Proc. Natl. Acad. Sci. USA117, 10625 (2020)
work page 2020
-
[7]
R. Schaeffer, M. Khona, Z. Robertson, A. Boopathy, K. Pistunova, J. W. Rocks, I. R. Fiete, and O. Koyejo (2023), 2303.14151
arXiv 2023
- [8]
Show all 21 references
-
[9]
Speicher and J
R. Speicher and J. Hoffmann,High-dimensional analysis: Random matrices and machine learning(2023), lecture notes, Saarland University, Summer Term 2023
2023
-
[10]
Penington, S
G. Penington, S. H. Shenker, D. Stanford, and Z. Yang, JHEP03, 205 (2022), 1911.11977
2022 arXiv
-
[11]
Kawabata, T
K. Kawabata, T. Nishioka, Y. Okuyama, and K. Watanabe, JHEP05, 062 (2021), 2102.02425
2021 arXiv
-
[12]
J. M. V. Balasubramanian, A. Lawrence and M. Sasieta, Phys. Rev. X14, 011024 (2024)
2024
-
[13]
J. M. M. V. Balasubramanian, A. Lawrence and M. Sasieta, Phys. Rev. Lett.132, 141501 (2024)
2024
-
[14]
J. M. M. M. S. A. Climent, R. Emparan and A. V. L´ opez, Phys. Rev. D109, 086024 (2024)
2024
- [15]
- [16]
-
[17]
Mehta, M
P. Mehta, M. Bukov, C. H. Wang, A. G. R. Day, C. Richardson, C. K. Fisher, and D. J. Schwab, Phys. Rep.810, 1 (2019)
2019
-
[18]
Kempkes, A
M. Kempkes, A. Ijaz, E. Gil-Fuster, C. Bravo-Prieto, J. Spiegelberg, E. van Nieuwenburg, and V. Dunjko (2025), 2501.10077
2025
- [19]
-
[20]
S. L. Braunstein and A. K. Pati, Phys. Rev. Lett.98, 080502 (2007), gr-qc/0603046
2007 arXiv
-
[21]
J. W. Rocks and P. Mehta, Physical Review Research4, 013201 (2022)
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.