Pith. sign in

REVIEW 2 major objections 5 minor 10 references

Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling

T0 review · 2 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Two-layer NTK spectrum is a Marchenko–Pastur map of an explicit law.

desk verdict The quadratic-scaling NTK spectrum is a real new result, but the abstract overclaims: the proof relies on a sub-Gaussian-like concentration inequality that the 'i.i.d. with four moments' framing doesn't provide. read the letter →

arxiv 2508.20036 v1 pith:5DTAGEUO submitted 2025-08-27 math.PR stat.ML

classification math.PRstat.ML MSC 60B2046L54
keywords neuraltangentkerneleigenvaluedistributionMarchenko–PasturmapfreemultiplicativeconvolutionquadraticscalingHadamardproductrandommatrixtheoryinterpolationthreshold
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves that in the quadratic scaling of a two-layer neural network—where the number of samples n, input dimension d, and hidden width p grow with n/(dp)→γ1 and p/d→γ2, the regime around the interpolation threshold—the eigenvalue distribution of the neural tangent kernel is asymptotically deterministic. The limit is the Marchenko–Pastur map with shape γ1 applied to a deterministic measure μ_{ν,φ} that depends only on the law of the squared readout weights, the width-aspect γ2, and the activation derivative φ; the full NTK and its Hadamard-product part share this bulk limit because the extra conjugate-kernel term has rank O(√n). When D²=I or φ is linear, the measure is fully explicit: μ^{γ1}_{MP}⊠Law(α²_φ χ+β²_ψ), where χ is a classical convolution of two Marchenko–Pastur laws plus an atom at zero, and αφ, βψ are the Gaussian first-order correlation and residual variance of φ. The proof identifies the NTK as a Gram matrix of tensor-product rows, applies the Bai–Zhou theorem for sample covariance matrices with independent rows, and then computes the spectrum of the resulting covariance tensor via its moments and free probability.

What carries the argument

The Marchenko–Pastur map μ^{γ}_{MP}⊠ν, whose Stieltjes transform solves s(z)=∫ (t(1−γ(1+zs(z)))−z)^{-1} dν(t), is the object carrying the result: it sends the spectrum of a population covariance to the spectrum of a sample covariance with sample-to-feature ratio γ. The proof's workhorse is the reduction of the kernel to a Gram matrix of vectors ω_i=x_i⊗D(αφ/√d W^T x_i+ψ(ẽx_i)) in Rd⊗Rp, so that Bai–Zhou's theorem applies with independent rows. The limiting covariance tensor Q factors as D bQ D, and bQ's spectrum is computable exactly from the singular values of W: eigenvalues α²φ/d(σ²_i+σ²_j)+β²ψ plus a β²ψ eigenspace, which yields the explicit Law(α²φχ+β²ψ) in the commuting cases. In genera

What would settle it

Run the explicit case ν=δ1, φ(x)=x at a new (γ1,γ2) pair, e.g. γ1=0.7, γ2=1.4, with n around 20,000, and compare the empirical density of K to μ^{γ1}_{MP}⊠Law(χ), Law(χ)=γ2/2 μ^{γ2}_{MP}∗μ^{γ2}_{MP}+(1−γ2)μ^{γ2}_{MP}+γ2/2 δ0; any systematic deviation that persists as n grows would refute the main formula. To probe the load-bearing hypothesis instead, use centered data with a finite fourth moment but row norms that concentrate only polynomially at rate d^{-δ}, δ<1/2; the theorem's predicted MP-map limit should fail, revealing that Assumption 2.2 is doing real work.

Watch

Extended reading notes

Core claim

The central discovery is a limit theorem: under Assumptions 2.1–2.4, the empirical spectral distribution of K=(1/d)XX^T⊙(1/p)φ(XW/√d)D²φ(XW/√d)^T, and of the full two-layer NTK, converges weakly in probability to μ^{γ1}_{MP}⊠μ_{ν,φ}. The measure μ_{ν,φ} is defined by moment formulas (Proposition 5.10) and by the free-probabilistic Stieltjes transform in Proposition 5.12. In the two cases where the shift is trivial, D²=I or φ linear, the law collapses to μ^{γ1}_{MP}⊠Law(α²_φ χ+β²_ψ) with Law(χ)=γ2/2(μ^{γ2}_{MP}⊠ν)∗(μ^{γ2}_{MP}⊠ν)+(1−γ2)(μ^{γ2}_{MP}⊠ν)+γ2/2 δ0; here αφ is the Gaussian covariance of φ with the identity, ψ is the residual after removing the constant and linear Gaussian chaoses,

Load-bearing premise

The proof requires every data row's squared norm to be within d^{-1/2+ε} of its expectation with failure probability decaying faster than any polynomial; this concentration is much stronger than the stated fourth-moment condition and effectively assumes sub-exponential (near-bounded) data entries.

Editorial extensions

If this is right

  • The eigenvalue bulk of the full NTK is identical to that of the Hadamard-only matrix K, so the conjugate-kernel term created by differentiating the outer weights does not shape the spectrum at this scaling.
  • For D²=I or linear activations, the limiting spectrum can be written down by convolving two Marchenko–Pastur densities with weights γ2/2 and 1−γ2 and then applying the Marchenko–Pastur map of shape γ1; no simulation is needed.
  • The limit depends on the activation only through αφ=E[zφ(z)] and the residual variance β²ψ=E[ψ(z)²], so many activations share the same bulk spectrum when these two numbers match.
  • In the general case the spectrum is characterized by the Stieltjes transform (5.4) in terms of the free non-commuting pair (D², (1/d)W^TW), giving a complete algorithm for the limit from moment data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The proof's reliance on Assumption 2.2 suggests a sharp-threshold question the paper leaves open: how many moments, or what concentration rate, are actually needed for the MP-map limit; a plausible answer is sub-Gaussian or bounded rows, with heavy-tailed rows producing a different limit.
  • The tensor-Gram rewrite is not specific to the two-layer NTK; the same strategy should give MP-map-of-covariance-tensor limits for Hadamard products of other nonlinear feature matrices, including deep linearizations, at the interpolation threshold.
  • Since the noncommuting shift is the only obstruction to the simple convolution formula, choosing D with non-constant entries and a nonlinear φ should produce spectra measurably different from the naive Law(α²φχ+β²ψ), which is a direct numerical test of Proposition 5.12.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper computes the asymptotic empirical spectral distribution of the neural tangent kernel matrix K = (1/d) X X^T ⊙ (1/p) φ(XW/√d) D^2 φ(XW/√d)^T in the quadratic scaling n/(dp)→γ1, p/d→γ2, with X i.i.d., W standard Gaussian, and D diagonal with bounded i.i.d. entries. The main result, Theorem 2.7, states that the empirical spectral distributions of K and of the full two-layer NTK converge weakly in probability to μ_{MP}^{γ1} ⊠ μ_{ν,φ}, where μ_{ν,φ} is characterized by a free-probability formula in Section 5.2. In the special cases D^2 = I or φ linear, the limit is explicitly μ_{MP}^{γ1} ⊠ Law(α_φ^2 χ + β_ψ^2), where Law(χ) is a classical convolution of Marchenko–Pastur-type measures. The proof strategy is: (1) replace the nonlinearity φ by a linear part plus independent Gaussian noise (Proposition 3.1); (2) rewrite the model as a Gram matrix of random tensors and apply the Bai–Zhou sample-covariance theorem (Section 4); (3) compute the limiting spectrum of the covariance tensor Q via exact finite-rank structures and moment/free-probability calculations (Section 5).

Significance. If the result is correct, it is a significant contribution to nonlinear random matrix theory and to the spectral analysis of neural tangent kernels in the interpolation regime. The paper contains no fitted parameters: γ1 and γ2 are dimension ratios and α_φ, β_ψ are fixed Gaussian inner products; the special cases in Corollary 2.8 are explicit and match the four simulation configurations in Figure 2. The general free-probability formula in Proposition 5.12 is a genuinely new structural description, and the observation that the first-order chaos produces a convolution structure while higher chaoses act like an additive shift is illuminating. The main caveats are two unstated or under-proved steps in the general proof: the reduction from the covariance tensor T to its low-rank perturbation term, and the strength of the concentration assumption relative to the advertised i.i.d. data setting. Both are fixable, but they need to be addressed before the paper can be accepted.

major comments (2)
  1. The transition from the covariance tensor T to the tensor Q is asserted, not proved. After Lemma 4.1, T = Q + R^(1) with R^(1) = (α_φ^2/d) μ_{4,x} diag1(WD ⊗ WD). Section 5 begins 'From the previous section, we know that the limiting eigenvalue distribution of eK is given by μ_MP^{γ1} ⊠ π where π is the asymptotic e.e.d of tensor Q, the leading part of tensor T.' No argument is given that replacing T by Q does not change the limiting empirical eigenvalue distribution required by the Bai–Zhou theorem (Theorem 4.2). As defined, diag1 is supported on q=s and R^(1) is, as an operator, a direct sum of rank-one blocks over the first index, hence has rank at most d while the tensor space has dimension dp; this rank estimate would make rank/(dp)=O(1/p) → 0 and would justify the reduction, but this is not stated. Without this explicit rank/closeness argument, Theorem 2.7 does not follow from Theo
  2. The concentration inequality (2.1) is substantially stronger than the stated 'i.i.d. centered random variables of variance 1 such that μ_{4,x} exists' hypothesis. It demands P(|‖x_i‖²/d − 1| ≥ d^ε/√d) ≤ N^{-θ} for every ε,θ > 0, which is effectively a sub-exponential/bounded-entry row concentration property. For i.i.d. entries with only four moments, Chebyshev gives only a polynomial bound of order d^{-2ε}, not N^{-θ} for all θ. Inequality (2.1) is used in the load-bearing estimates of Lemma 3.5, Lemma 3.6, Section 4, and Theorem A.1. Thus the theorem as proven is narrower than the abstract's 'X ∈ R^{n×d} is an i.i.d random matrix' and the introduction's claim that Gaussian data is not needed. The authors should either state the strong tail assumption prominently in the abstract and Section 1, or replace it with an explicit bounded/sub-Gaussian condition and adapt the proofs.
minor comments (5)
  1. [Assumption 2.2] The parameter N in (2.1) is never defined; it should be n, d, or another growth parameter. This makes the assumption ambiguous.
  2. [Section 2, Theorem 2.6] There is a typo: 'Suppose that A ⪰ 0 and B ⪰ are a × a and b × b matrices' should read 'B ⪰ 0'.
  3. [Section 3, Lemma 3.5 proof] The proof concludes with O(d^{4ε}p) while the statement claims C p d^ε. This is not an error, but the epsilon bookkeeping should be made uniform so the reader can verify the claimed rate for all ε in the statement.
  4. [Section 3, Proposition 3.1] The proof sums the individual high-probability estimates of Lemma 3.4 over i without explicitly discussing the union bound over i = 1,...,n or the correlation between the summands. A sentence explaining the union bound and the deterministic factor 1/n would help.
  5. [Title and Section 1] The title contains a typo ('Eigenv alue') and Section 1 contains 'linear regine'; these should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the claimed limit is derived from the model via standard random matrix results, with no fitted parameter or self-citation forcing the conclusion.

full rationale

The paper's central claim (Theorem 2.7) is that the empirical spectral distribution of the NTK matrix converges to the Marchenko–Pastur map applied to a deterministic measure. That measure is not assumed or fitted: it is computed from the model. The parameters γ1 and γ2 are limits of dimension ratios; α_φ and β_ψ are Gaussian moments of the fixed activation φ; ν is the law of a². None of these is tuned to match the target eigenvalue distribution. The proof chain is self-contained: Proposition 3.1 replaces the nonlinearity by a Gaussian counterpart using concentration estimates proved in the paper; Section 4 verifies the Bai–Zhou Gram-matrix theorem; Section 5 computes the covariance tensor's limiting spectrum and moments, leading to the free-probability Stieltjes transform formula in Proposition 5.12. The special cases in Corollary 2.8 follow from explicit spectral computations (Lemma 5.3) and the Marchenko–Pastur limit theorem, not from assuming the result. Citations to prior work, including [BZ08], [BS10], and even self-citations such as [Paq24], are used for standard technical lemmas (e.g., sample-covariance resolvent bounds) whose assumptions do not include the target result; they are not used to define away the conclusion. The only notable weakness is that Assumption 2.2's concentration inequality (2.1) is much stronger than the stated fourth-moment condition, so the abstract's phrase 'i.i.d random matrix' is broader than the theorem actually requires. This is a hypothesis-strength / scope concern, not a circularity: the derivation still computes the limit from the model rather than importing it as an input.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on standard random matrix theorems and on the stated distributional assumptions. No free parameters are fitted: gamma1 and gamma2 are limits of deterministic aspect ratios, and alpha_phi, beta_psi are Gaussian inner products of the fixed activation phi. The only potentially fragile input is the concentration inequality (2.1), which is an explicit assumption but is stronger than the fourth-moment hypothesis it accompanies.

assumptions (5)
  • standard math Bai-Zhou theorem (Theorem 4.2, [BZ08]) for sample covariance matrices with independent but non-identically distributed rows
    Used to assert the Marchenko-Pastur map for the Gram matrix of the eomega_i rows, conditional on W and D, given convergence of the covariance tensor; Section 4.
  • standard math Marchenko-Pastur map and Bai-Silverstein theorems [BS10, Theorem 2.6]
    Used to define the MP map and for rank-perturbation bounds on Stieltjes transforms; Sections 2 and 3.
  • domain assumption Asymptotic freeness of (D^2, (1/d)W^TW) in Proposition 5.12
    The general characterization of mu_{nu,phi} assumes this pair converges in non-commutative moments to free variables; standard for independent Gaussian W and diagonal D, but stated as a hypothesis.
  • domain assumption Concentration inequality (2.1) for data row norms with arbitrary polynomial tails
    Assumed in Assumption 2.2 and used pervasively (Lemma 3.5, Lemma 3.6, Section 4, Theorem A.1); effectively requires sub-exponential data rather than just finite fourth moment.
  • domain assumption Pseudo-Lipschitz phi with Gaussian chaos decomposition
    Assumption 2.4; enables the orthogonal decomposition phi(x) = c + alpha x + psi(x), which is the starting point of the proof in Section 3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling." pith.science (2026). https://pith.science/paper/5DTAGEUO

@misc{pith2026250820036,
  author       = {Pith},
  title        = {Pith review of: Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5DTAGEUO}},
  note         = {Machine review of arXiv:2508.20036}
}
abstract

We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times p}$ is an i.i.d $\mathcal{N}(0,1)$ matrix and $D\in\mathbb{R}^{p\times p}$ is a diagonal matrix with i.i.d bounded entries, we consider the matrix \[ \mathrm{NTK} = \frac{1}{d}XX^\top \odot \frac{1}{p} \sigma'\left( \frac{1}{\sqrt{d}}XW \right)D^2 \sigma'\left( \frac{1}{\sqrt{d}}XW \right)^\top \] where $\sigma'$ is a pseudo-Lipschitz function applied entrywise and under the scaling $\frac{n}{dp}\to \gamma_1$ and $\frac{p}{d}\to \gamma_2$. We describe the asymptotic distribution as the free multiplicative convolution of the Marchenko--Pastur distribution with a deterministic distribution depending on $\sigma$ and $D$.

Figures

Figures reproduced from arXiv: 2508.20036 by the authors.

Figure 1
Figure 1. Illustration of a multi-layered feed-forward neural network [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Eigenvalue density comparisons for the model ( [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Diagrammatic representation of the matching [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Illustration of the procedure to bound the example term above, we start by drawing the graph [PITH_FULL_IMAGE:figures/full_fig_p032_4.png]
Figure 5
Figure 5. Figure 5: After replacing the edge {ui−1, ui+1} to {ui−1, ui} and {ui, ui+1} doing the same for the red edge {ui, ui+1}, we see that ui is of degree 4. the procedure of ui is stil 4 since this edge is replaced by {ui, ui+1} and {ui+1, ui+2} and thus only one is incident to ui. T…
Figure 6
Figure 6. Figure 6: If a group of adjacent vertices are identified, the corresponding vertex has some self-loops and [PITH_FULL_IMAGE:figures/full_fig_p033_6.png]
Figure 7
Figure 7. Figure 7: Example of a general procedure. In the first step, we replace the red edges by the two corre [PITH_FULL_IMAGE:figures/full_fig_p033_7.png]
Figure 8
Figure 8. Figure 8: Example of a term which gives rise to a graph with a single vertex. [PITH_FULL_IMAGE:figures/full_fig_p034_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

10 extracted references · 5 canonical work pages

  1. [1]

    Abou Assaly and L

    [AAB25] S. Abou Assaly and L. Benigni,Eigenvalue distribution of the hadamard product of sample covariance matrices in a quadratic regime, arXiv e-prints (2025), arXiv–2502. [AP20] B. Adlam and J. Pennington,The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization, International conference on machine learning...

  2. [949]

    Extremal Eigenvalues of Random Kernel Matrices with Polynomial Scaling

    [JGH18] A. Jacot, F. Gabriel, and C. Hongler,Neural tangent kernel: convergence and generalization in neural networks, Proceedings of the 32nd international conference on neural information processing systems, 2018, pp. 8580–8589. [Kar10] N. El Karoui,The spectrum of kernel random matrices, The Annals of Statistics38 (2010), no. 1, 1 –50. [KNH24] D. Kogan...

  3. [2010]

    Probab.26 (1998), no

    [BS98] , No eigenvalues outside the support of the limiting spectral distribution of large-dimensional sample covariance matrices, Ann. Probab.26 (1998), no. 1, 316 –345. [BZ08] Z. Bai and W. Zhou,Large sample covariance matrices without independence structures in columns, Statist. Sinica 18 (2008), no. 2, 425–442. [CBP21] A. Canatar, B. Bordelon, and C. ...

  4. [2012]

    Louart, Z

    [LLC18] C. Louart, Z. Liao, and R. Couillet,A random matrix approach to neural networks, The Annals of Applied Probability 28 (2018), no. 2, 1190–1248. [LM21] Z. Liao and M. W. Mahoney,Hessian eigenspectra of more realistic nonlinear models, Advances in neural information processing systems, 2021, pp. 20104–20117. [LRZ20] T. Liang, A. Rakhlin, and X. Zhai...

  5. [2017]

    Montanari and Y

    [MZ22] A. Montanari and Y. Zhong,The interpolation phase transition in neural networks: Memorization and generalization under lazy training, The Annals of Statistics50 (2022), no. 5, 2816–2847. [NTS15] B. Neyshabur, R. Tomioka, and N. Srebro,In search of the real inductive bias: On the role of implicit regularization in deep learning, 3rd international co...

  6. [2019]

    [Fel25] M. J. Feldman,Spectral properties of elementwise-transformed spiked matrices, SIAM Journal on Mathematics of Data Science7 (2025), no. 2, 542–571. [FM19] Z. Fan andA. Montanari,The spectral norm of random inner-product kernel matrices, Probability Theory andRelated Fields 173 (2019), no. 1, 27–85. [FM24] Z. Fan and R. Ma, Kronecker-product random ...

  7. [2020]

    [HLM24] H. Hu, Y. M. Lu, and T. Misiakiewicz,Asymptotics of random feature regression beyond the linear scaling regime, arXiv preprint arXiv:2403.08160 (2024). [HMRT22] T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani,Surprises in high-dimensional ridgeless least squares inter- polation, Annals of statistics50 (2022), no. 2,

  8. [2022]

    Chizat, E

    [COB19] L. Chizat, E. Oyallon, and F. Bach,On lazy training in differentiable programming, Advances in neural information processing systems32 (2019). [CS13] X. Cheng and A. Singer,The spectrum of random inner-product kernel matrices, Random Matrices: Theory and Applications 02 (2013), no. 04, 1350010. [DCLT19] J. Devlin, M.-W. Chang, K. Lee, and K. Touta...

Show all 10 references
  1. [2024]

    Piccolo and D

    [PS21] V. Piccolo and D. Schröder, Analysis of one-hidden-layer neural networks via the resolvent method, Advances in neural information processing systems34 (2021), 5225–5235. [PW17] J. Pennington and P. Worah,Nonlinear random matrix theory for deep learning, Advances in neur...

  2. [2914]

    Chouard, Deterministic equivalent of the conjugate kernel matrix associated to artificial neural networks, arXiv preprint arXiv:2306.05850 (2023)

    [Cho23] C. Chouard, Deterministic equivalent of the conjugate kernel matrix associated to artificial neural networks, arXiv preprint arXiv:2306.05850 (2023). [CL22] R. Couillet and Z. Liao,Random matrix methods for machine learning, Cambridge University Press,

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.