REVIEW 2 major objections 5 minor 10 references
Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
T0 review · 2 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Two-layer NTK spectrum is a Marchenko–Pastur map of an explicit law.
desk verdict The quadratic-scaling NTK spectrum is a real new result, but the abstract overclaims: the proof relies on a sub-Gaussian-like concentration inequality that the 'i.i.d. with four moments' framing doesn't provide. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Marchenko–Pastur map μ^{γ}_{MP}⊠ν, whose Stieltjes transform solves s(z)=∫ (t(1−γ(1+zs(z)))−z)^{-1} dν(t), is the object carrying the result: it sends the spectrum of a population covariance to the spectrum of a sample covariance with sample-to-feature ratio γ. The proof's workhorse is the reduction of the kernel to a Gram matrix of vectors ω_i=x_i⊗D(αφ/√d W^T x_i+ψ(ẽx_i)) in Rd⊗Rp, so that Bai–Zhou's theorem applies with independent rows. The limiting covariance tensor Q factors as D bQ D, and bQ's spectrum is computable exactly from the singular values of W: eigenvalues α²φ/d(σ²_i+σ²_j)+β²ψ plus a β²ψ eigenspace, which yields the explicit Law(α²φχ+β²ψ) in the commuting cases. In genera
What would settle it
Run the explicit case ν=δ1, φ(x)=x at a new (γ1,γ2) pair, e.g. γ1=0.7, γ2=1.4, with n around 20,000, and compare the empirical density of K to μ^{γ1}_{MP}⊠Law(χ), Law(χ)=γ2/2 μ^{γ2}_{MP}∗μ^{γ2}_{MP}+(1−γ2)μ^{γ2}_{MP}+γ2/2 δ0; any systematic deviation that persists as n grows would refute the main formula. To probe the load-bearing hypothesis instead, use centered data with a finite fourth moment but row norms that concentrate only polynomially at rate d^{-δ}, δ<1/2; the theorem's predicted MP-map limit should fail, revealing that Assumption 2.2 is doing real work.
Extended reading notes
Core claim
The central discovery is a limit theorem: under Assumptions 2.1–2.4, the empirical spectral distribution of K=(1/d)XX^T⊙(1/p)φ(XW/√d)D²φ(XW/√d)^T, and of the full two-layer NTK, converges weakly in probability to μ^{γ1}_{MP}⊠μ_{ν,φ}. The measure μ_{ν,φ} is defined by moment formulas (Proposition 5.10) and by the free-probabilistic Stieltjes transform in Proposition 5.12. In the two cases where the shift is trivial, D²=I or φ linear, the law collapses to μ^{γ1}_{MP}⊠Law(α²_φ χ+β²_ψ) with Law(χ)=γ2/2(μ^{γ2}_{MP}⊠ν)∗(μ^{γ2}_{MP}⊠ν)+(1−γ2)(μ^{γ2}_{MP}⊠ν)+γ2/2 δ0; here αφ is the Gaussian covariance of φ with the identity, ψ is the residual after removing the constant and linear Gaussian chaoses,
Load-bearing premise
The proof requires every data row's squared norm to be within d^{-1/2+ε} of its expectation with failure probability decaying faster than any polynomial; this concentration is much stronger than the stated fourth-moment condition and effectively assumes sub-exponential (near-bounded) data entries.
Editorial extensions
If this is right
- The eigenvalue bulk of the full NTK is identical to that of the Hadamard-only matrix K, so the conjugate-kernel term created by differentiating the outer weights does not shape the spectrum at this scaling.
- For D²=I or linear activations, the limiting spectrum can be written down by convolving two Marchenko–Pastur densities with weights γ2/2 and 1−γ2 and then applying the Marchenko–Pastur map of shape γ1; no simulation is needed.
- The limit depends on the activation only through αφ=E[zφ(z)] and the residual variance β²ψ=E[ψ(z)²], so many activations share the same bulk spectrum when these two numbers match.
- In the general case the spectrum is characterized by the Stieltjes transform (5.4) in terms of the free non-commuting pair (D², (1/d)W^TW), giving a complete algorithm for the limit from moment data.
Reading between the lines
- The proof's reliance on Assumption 2.2 suggests a sharp-threshold question the paper leaves open: how many moments, or what concentration rate, are actually needed for the MP-map limit; a plausible answer is sub-Gaussian or bounded rows, with heavy-tailed rows producing a different limit.
- The tensor-Gram rewrite is not specific to the two-layer NTK; the same strategy should give MP-map-of-covariance-tensor limits for Hadamard products of other nonlinear feature matrices, including deep linearizations, at the interpolation threshold.
- Since the noncommuting shift is the only obstruction to the simple convolution formula, choosing D with non-constant entries and a nonlinear φ should produce spectra measurably different from the naive Law(α²φχ+β²ψ), which is a direct numerical test of Proposition 5.12.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper computes the asymptotic empirical spectral distribution of the neural tangent kernel matrix K = (1/d) X X^T ⊙ (1/p) φ(XW/√d) D^2 φ(XW/√d)^T in the quadratic scaling n/(dp)→γ1, p/d→γ2, with X i.i.d., W standard Gaussian, and D diagonal with bounded i.i.d. entries. The main result, Theorem 2.7, states that the empirical spectral distributions of K and of the full two-layer NTK converge weakly in probability to μ_{MP}^{γ1} ⊠ μ_{ν,φ}, where μ_{ν,φ} is characterized by a free-probability formula in Section 5.2. In the special cases D^2 = I or φ linear, the limit is explicitly μ_{MP}^{γ1} ⊠ Law(α_φ^2 χ + β_ψ^2), where Law(χ) is a classical convolution of Marchenko–Pastur-type measures. The proof strategy is: (1) replace the nonlinearity φ by a linear part plus independent Gaussian noise (Proposition 3.1); (2) rewrite the model as a Gram matrix of random tensors and apply the Bai–Zhou sample-covariance theorem (Section 4); (3) compute the limiting spectrum of the covariance tensor Q via exact finite-rank structures and moment/free-probability calculations (Section 5).
Significance. If the result is correct, it is a significant contribution to nonlinear random matrix theory and to the spectral analysis of neural tangent kernels in the interpolation regime. The paper contains no fitted parameters: γ1 and γ2 are dimension ratios and α_φ, β_ψ are fixed Gaussian inner products; the special cases in Corollary 2.8 are explicit and match the four simulation configurations in Figure 2. The general free-probability formula in Proposition 5.12 is a genuinely new structural description, and the observation that the first-order chaos produces a convolution structure while higher chaoses act like an additive shift is illuminating. The main caveats are two unstated or under-proved steps in the general proof: the reduction from the covariance tensor T to its low-rank perturbation term, and the strength of the concentration assumption relative to the advertised i.i.d. data setting. Both are fixable, but they need to be addressed before the paper can be accepted.
major comments (2)
- The transition from the covariance tensor T to the tensor Q is asserted, not proved. After Lemma 4.1, T = Q + R^(1) with R^(1) = (α_φ^2/d) μ_{4,x} diag1(WD ⊗ WD). Section 5 begins 'From the previous section, we know that the limiting eigenvalue distribution of eK is given by μ_MP^{γ1} ⊠ π where π is the asymptotic e.e.d of tensor Q, the leading part of tensor T.' No argument is given that replacing T by Q does not change the limiting empirical eigenvalue distribution required by the Bai–Zhou theorem (Theorem 4.2). As defined, diag1 is supported on q=s and R^(1) is, as an operator, a direct sum of rank-one blocks over the first index, hence has rank at most d while the tensor space has dimension dp; this rank estimate would make rank/(dp)=O(1/p) → 0 and would justify the reduction, but this is not stated. Without this explicit rank/closeness argument, Theorem 2.7 does not follow from Theo
- The concentration inequality (2.1) is substantially stronger than the stated 'i.i.d. centered random variables of variance 1 such that μ_{4,x} exists' hypothesis. It demands P(|‖x_i‖²/d − 1| ≥ d^ε/√d) ≤ N^{-θ} for every ε,θ > 0, which is effectively a sub-exponential/bounded-entry row concentration property. For i.i.d. entries with only four moments, Chebyshev gives only a polynomial bound of order d^{-2ε}, not N^{-θ} for all θ. Inequality (2.1) is used in the load-bearing estimates of Lemma 3.5, Lemma 3.6, Section 4, and Theorem A.1. Thus the theorem as proven is narrower than the abstract's 'X ∈ R^{n×d} is an i.i.d random matrix' and the introduction's claim that Gaussian data is not needed. The authors should either state the strong tail assumption prominently in the abstract and Section 1, or replace it with an explicit bounded/sub-Gaussian condition and adapt the proofs.
minor comments (5)
- [Assumption 2.2] The parameter N in (2.1) is never defined; it should be n, d, or another growth parameter. This makes the assumption ambiguous.
- [Section 2, Theorem 2.6] There is a typo: 'Suppose that A ⪰ 0 and B ⪰ are a × a and b × b matrices' should read 'B ⪰ 0'.
- [Section 3, Lemma 3.5 proof] The proof concludes with O(d^{4ε}p) while the statement claims C p d^ε. This is not an error, but the epsilon bookkeeping should be made uniform so the reader can verify the claimed rate for all ε in the statement.
- [Section 3, Proposition 3.1] The proof sums the individual high-probability estimates of Lemma 3.4 over i without explicitly discussing the union bound over i = 1,...,n or the correlation between the summands. A sentence explaining the union bound and the deterministic factor 1/n would help.
- [Title and Section 1] The title contains a typo ('Eigenv alue') and Section 1 contains 'linear regine'; these should be corrected.
Circularity Check
No significant circularity: the claimed limit is derived from the model via standard random matrix results, with no fitted parameter or self-citation forcing the conclusion.
full rationale
The paper's central claim (Theorem 2.7) is that the empirical spectral distribution of the NTK matrix converges to the Marchenko–Pastur map applied to a deterministic measure. That measure is not assumed or fitted: it is computed from the model. The parameters γ1 and γ2 are limits of dimension ratios; α_φ and β_ψ are Gaussian moments of the fixed activation φ; ν is the law of a². None of these is tuned to match the target eigenvalue distribution. The proof chain is self-contained: Proposition 3.1 replaces the nonlinearity by a Gaussian counterpart using concentration estimates proved in the paper; Section 4 verifies the Bai–Zhou Gram-matrix theorem; Section 5 computes the covariance tensor's limiting spectrum and moments, leading to the free-probability Stieltjes transform formula in Proposition 5.12. The special cases in Corollary 2.8 follow from explicit spectral computations (Lemma 5.3) and the Marchenko–Pastur limit theorem, not from assuming the result. Citations to prior work, including [BZ08], [BS10], and even self-citations such as [Paq24], are used for standard technical lemmas (e.g., sample-covariance resolvent bounds) whose assumptions do not include the target result; they are not used to define away the conclusion. The only notable weakness is that Assumption 2.2's concentration inequality (2.1) is much stronger than the stated fourth-moment condition, so the abstract's phrase 'i.i.d random matrix' is broader than the theorem actually requires. This is a hypothesis-strength / scope concern, not a circularity: the derivation still computes the limit from the model rather than importing it as an input.
Assumptions & free parameters
assumptions (5)
- standard math Bai-Zhou theorem (Theorem 4.2, [BZ08]) for sample covariance matrices with independent but non-identically distributed rows
- standard math Marchenko-Pastur map and Bai-Silverstein theorems [BS10, Theorem 2.6]
- domain assumption Asymptotic freeness of (D^2, (1/d)W^TW) in Proposition 5.12
- domain assumption Concentration inequality (2.1) for data row norms with arbitrary polynomial tails
- domain assumption Pseudo-Lipschitz phi with Gaussian chaos decomposition
Cite this review
Pith. "Pith review of Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling." pith.science (2026). https://pith.science/paper/5DTAGEUO
@misc{pith2026250820036,
author = {Pith},
title = {Pith review of: Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling},
year = {2026},
howpublished = {\url{https://pith.science/paper/5DTAGEUO}},
note = {Machine review of arXiv:2508.20036}
}
abstract
We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times p}$ is an i.i.d $\mathcal{N}(0,1)$ matrix and $D\in\mathbb{R}^{p\times p}$ is a diagonal matrix with i.i.d bounded entries, we consider the matrix \[ \mathrm{NTK} = \frac{1}{d}XX^\top \odot \frac{1}{p} \sigma'\left( \frac{1}{\sqrt{d}}XW \right)D^2 \sigma'\left( \frac{1}{\sqrt{d}}XW \right)^\top \] where $\sigma'$ is a pseudo-Lipschitz function applied entrywise and under the scaling $\frac{n}{dp}\to \gamma_1$ and $\frac{p}{d}\to \gamma_2$. We describe the asymptotic distribution as the free multiplicative convolution of the Marchenko--Pastur distribution with a deterministic distribution depending on $\sigma$ and $D$.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
[AAB25] S. Abou Assaly and L. Benigni,Eigenvalue distribution of the hadamard product of sample covariance matrices in a quadratic regime, arXiv e-prints (2025), arXiv–2502. [AP20] B. Adlam and J. Pennington,The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization, International conference on machine learning...
arXiv 2025
-
[949]
Extremal Eigenvalues of Random Kernel Matrices with Polynomial Scaling
[JGH18] A. Jacot, F. Gabriel, and C. Hongler,Neural tangent kernel: convergence and generalization in neural networks, Proceedings of the 32nd international conference on neural information processing systems, 2018, pp. 8580–8589. [Kar10] N. El Karoui,The spectrum of kernel random matrices, The Annals of Statistics38 (2010), no. 1, 1 –50. [KNH24] D. Kogan...
work page Pith review arXiv 2018
-
[2010]
[BS98] , No eigenvalues outside the support of the limiting spectral distribution of large-dimensional sample covariance matrices, Ann. Probab.26 (1998), no. 1, 316 –345. [BZ08] Z. Bai and W. Zhou,Large sample covariance matrices without independence structures in columns, Statist. Sinica 18 (2008), no. 2, 425–442. [CBP21] A. Canatar, B. Bordelon, and C. ...
work page 1998
-
[2012]
[LLC18] C. Louart, Z. Liao, and R. Couillet,A random matrix approach to neural networks, The Annals of Applied Probability 28 (2018), no. 2, 1190–1248. [LM21] Z. Liao and M. W. Mahoney,Hessian eigenspectra of more realistic nonlinear models, Advances in neural information processing systems, 2021, pp. 20104–20117. [LRZ20] T. Liang, A. Rakhlin, and X. Zhai...
arXiv 2018
-
[2017]
[MZ22] A. Montanari and Y. Zhong,The interpolation phase transition in neural networks: Memorization and generalization under lazy training, The Annals of Statistics50 (2022), no. 5, 2816–2847. [NTS15] B. Neyshabur, R. Tomioka, and N. Srebro,In search of the real inductive bias: On the role of implicit regularization in deep learning, 3rd international co...
work page 2022
-
[2019]
[Fel25] M. J. Feldman,Spectral properties of elementwise-transformed spiked matrices, SIAM Journal on Mathematics of Data Science7 (2025), no. 2, 542–571. [FM19] Z. Fan andA. Montanari,The spectral norm of random inner-product kernel matrices, Probability Theory andRelated Fields 173 (2019), no. 1, 27–85. [FM24] Z. Fan and R. Ma, Kronecker-product random ...
-
[2020]
[HLM24] H. Hu, Y. M. Lu, and T. Misiakiewicz,Asymptotics of random feature regression beyond the linear scaling regime, arXiv preprint arXiv:2403.08160 (2024). [HMRT22] T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani,Surprises in high-dimensional ridgeless least squares inter- polation, Annals of statistics50 (2022), no. 2,
arXiv 2024
-
[2022]
[COB19] L. Chizat, E. Oyallon, and F. Bach,On lazy training in differentiable programming, Advances in neural information processing systems32 (2019). [CS13] X. Cheng and A. Singer,The spectrum of random inner-product kernel matrices, Random Matrices: Theory and Applications 02 (2013), no. 04, 1350010. [DCLT19] J. Devlin, M.-W. Chang, K. Lee, and K. Touta...
arXiv 2019
Show all 10 references
-
[2024]
Piccolo and D
[PS21] V. Piccolo and D. Schröder, Analysis of one-hidden-layer neural networks via the resolvent method, Advances in neural information processing systems34 (2021), 5225–5235. [PW17] J. Pennington and P. Worah,Nonlinear random matrix theory for deep learning, Advances in neur...
2021
-
[2914]
Chouard, Deterministic equivalent of the conjugate kernel matrix associated to artificial neural networks, arXiv preprint arXiv:2306.05850 (2023)
[Cho23] C. Chouard, Deterministic equivalent of the conjugate kernel matrix associated to artificial neural networks, arXiv preprint arXiv:2306.05850 (2023). [CL22] R. Couillet and Z. Liao,Random matrix methods for machine learning, Cambridge University Press,
2023 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.