REVIEW 3 major objections 6 minor 1 cited by
Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
T0 review · 3 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read In the large-N limit, kernel-regression training splits into N independent effective modes, and the test error is exactly one of the order parameters.
desk verdict A credible DMFT unifying Bayesian, gradient-flow, and Langevin regression with power-law spectra; main caveats are the unproven VGA exactness and a missing promised early-stopping relation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dynamical mean-field reduction: after disorder-averaging the stochastic training dynamics, the partition function is approximated by the best Gaussian process (variational Gaussian approximation), characterized by three order parameters: the response R(t,s), the correlation C(t,s), and the conjugate correlation C̃(t,s). These order parameters self-consistently determine a memory kernel K = (1−R)^{-1} and a colored noise with correlator 2β^{-1}δ + P η_i K*C̄*K^T. The resulting effective equation per mode i is (∂t + 1/(gβ))v_i + P η_i ∫K v_i = w̄_i/(gβ) + ξ_i, so the whole collective dynamics is encoded in how the eigenvalue η_i enters this single-mode equation.
What would settle it
Compute, from simulations at fixed finite N and P, the connected third- or fourth-order cumulant of the order parameters C(t,s) and R(t,s) over disorder realizations. If these cumulants are not small compared with the Gaussian predictions, the central-limit justification fails and the theory's error curves lose quantitative control. A sharper test: measure the actual noise correlator of the effective single-mode equations and compare it with the predicted P η_i K*C̄*K^T; a mismatch at accessible N directly falsifies the self-consistency.
Extended reading notes
Core claim
The paper establishes a dynamical mean-field theory for the typical learning dynamics of random-feature regression with power-law-distributed kernel eigenvalues. Its core result is a set of N decoupled effective equations, one for each eigenmode of the feature kernel, whose only mutual coupling is through a collective memory kernel K(t−s) and a common, self-consistently determined, time-correlated noise. The stationary point of the variational Gaussian approximation yields explicit closed equations for the mode means and Green's functions, and the disorder-averaged autocorrelation C̄(t,s) satisfies a self-consistent equation whose diagonal is the test error. The theory therefore yields not j
Load-bearing premise
Everything quantitative rests on the claim that, for large N, the variational Gaussian approximation becomes exact because the disorder-averaged dynamics are Gaussian; no error bound is given, and if non-Gaussianity is still sizable at the finite N and P used in practice, the theory's predicted error curves are uncontrolled.
Editorial extensions
If this is right
- Scaling laws in this model are not free-fitting curves: the power-law exponent γ of the kernel spectrum, the regularization strength, and the stopping time determine the generalization-error dynamics through one self-consistent equation.
- Early stopping emerges from the bias–variance trade-off in time: the bias falls monotonically as fast modes learn, while a self-reinforcing effective noise steadily grows the variance; the optimal stopping time is where the two cross.
- Bayesian GP inference (t→∞), gradient flow (β→∞), and finite-temperature Langevin dynamics are limits of one effective single-mode equation, so results transfer between these training regimes.
- Stronger L2 regularization has little effect on the location or depth of the early-stopping minimum, but it suppresses the late-time variance divergence, so it primarily helps when training noise is significant.
- Because large eigenmodes are learned earlier and more accurately, the model gives a dynamical account of spectral bias: the order in which modes are learned is governed by η_i, with small-eigenvalue modes slowing down all others through the collective kernel.
Reading between the lines
- Editorial inference: the self-consistency relation between the memory kernel and the noise means a single measured trajectory of one mode's variance, combined with the known spectrum η_i, pins down the collective kernel K; this gives a practical way to verify the theory on real training runs without averaging over many datasets.
- Editorial inference: if the variational Gaussian step has a slow, power-law approach to Gaussianity, then extrapolating scaling laws from small models to large ones—the practical motivation in the introduction—would inherit the same exponent-dependent error; measuring the fourth-order cumulant of C̄ at finite N would quantify how reliable such extrapolation is.
- Editorial inference: the same DMFT structure should apply beyond Gaussian feature statistics to any rotationally invariant random feature ensemble in which the empirical kernel is Wishart-like, so the theory's predictions for scaling-law exponents may be more universal than the specific Gaussian-feature model used to derive them.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives a dynamical mean-field theory for random-feature regression with power-law kernel eigenvalue spectra, starting from the MSRDJ path-integral representation and a variational Gaussian approximation (VGA). The effective theory reduces the disorder-averaged training dynamics to N independent effective Langevin equations, coupled only through a collective memory kernel K and a self-consistently determined noise correlator built from order parameters R, C, and C̃. The test error is expressed as 1/2 C̄(t,t). The framework is claimed to unify Bayesian inference, gradient flow, and Langevin training, and the authors compare the resulting predictions with simulations at N=P=100 for spectral bias, the bias-variance trade-off, regularization effects, and early stopping.
Significance. If the large-N assumptions are valid, the paper provides a compact, parameter-free reduction of high-dimensional random-feature learning dynamics to scalar order parameters, unifying prior static and deterministic treatments (Canatar et al., Advani et al., Maloney et al., Bordelon et al.) and giving a mechanistic explanation of early stopping. The derivation in Appendices A–C is detailed, and the simulations in Figures 1–4 serve as independent consistency checks with no fitted parameters; these are genuine strengths. The main reservations are that one advertised analytical result is absent and that the VGA/self-averaging argument lacks finite-N control, so the quantitative claim is not yet fully supported.
major comments (3)
- [Section 1 (Introduction), Sections 4–5] The Introduction promises "analytical results that relate the power-law exponent of the feature kernel, regularization, and early stopping time to obtain a minimal generalization gap." I could not find such a result in Sections 4–5. The paper provides a self-consistent integral-equation description and numerical evidence of early-stopping minima (Figs. 2–4), but no closed-form or asymptotic expression connects the exponent γ, the regularization gβ, and the stopping time to the minimal test error. Either supply the derivation (even in a special limit) or revise the advertised contribution; as written, one of the three central claims is not delivered.
- [Appendix C; Section 5; Eq. (35)] The derivation of the effective equations (19)–(25) rests on the variational Gaussian approximation applied to the action (35), whose nonlinear ln det term is non-quadratic. The only justification for exactness is the sentence in Section 5 that "the central limit theorem guarantees the Gaussianity of the process" as N→∞. No finite-N error bound or CLT statement (what is summed, in what sense) is given, and the term is not VGA-exact at finite N. All simulations use N=P=100 (Figs. 2–4) without error bars or a second value of N, so the agreement cannot separate a correct N→∞ theory from one with O(N^{-1/2}) bias. Please add finite-N scaling (e.g., N=200, 400 at fixed P/N), error bars, and a more precise large-N argument.
- [Section 4.1 and Section 5] The manuscript states that "the training process is indeed self-averaging" and that the order parameters R, C, C̃ "concentrate around their expectation values as N→∞," but no evidence or variance calculation is provided. Since the goal is typical-case behavior, the relevance of ⟨Z⟩_Q rather than ⟨ln Z⟩_Q depends on this concentration. The simulations average over 10^5 disorder realizations, but no single-realization trajectory or standard deviation of L_test is shown. A demonstration of vanishing fluctuations with N is needed to support the self-averaging claim.
minor comments (6)
- [Section 4.3, Eq. (29)] The expression L_test = 1/2 C̄(t,s) should be written as 1/2 C̄(t,t) (or with an explicit evaluation at s=t).
- [Eqs. (14), (19), (25)] The teacher weight ar{w}_i is sometimes typeset as w_i in the equations, which makes the teacher-student notation ambiguous. Please ensure overbars are consistently used.
- [Section 4.2, Eqs. (20)–(24)] K is first defined as a two-time kernel K(t,s), but it is then written as K(t−s) and solved in the Fourier domain. Clarify the causal/one-sided Fourier convention and how the t=0 initial condition is encoded in the Fourier solution.
- [Figure captions, Figs. 2–4] Please define η_1 and the normalization of Λ. The caption states Λ_ij = i^{-3/2} δ_ij, which corresponds to γ=1/2 in Eq. (9), but the value of η_1 is not given.
- [References] The reference entry for Naveh et al. is malformed ("journal = arXiv"). Please correct it.
- [Title and Abstract] The phrase "neural scaling laws" is used broadly. The numerical experiments use a single value of P and N, so the paper does not directly demonstrate L_test ∼ N^{-α} or P^{-α} scaling. Please either add experiments varying P and N or clarify in the abstract that the paper studies the dynamics that underlie such scaling laws without extracting the scaling exponents here.
Circularity Check
No significant circularity: order parameters are solved self-consistently; the test error is a defined observable, not a fitted input.
full rationale
The central claim — that L_test(t) = 1/2 C̄(t,t) (Eq. 29) follows from the dynamical mean-field theory — is not circular. Equation (29) is obtained from the definition of the test error in Eq. (7): averaging (y* − f*)² over the Gaussian feature distribution gives Σ_i η_i ⟨v_i²⟩, which is exactly the order parameter C̄(t,t) defined in (18). The value of C̄ is not assumed; it is the solution of the self-consistent system (19)–(28), derived from the disorder-averaged MSRDJ action (35) by a variational Gaussian approximation (Appendix C). No parameter is fitted to the simulation curves: Figs. 1–4 use the same model spectrum Λ_ij = i^(−3/2) and compare parameter-free theory solutions with independent simulations. The VGA's claimed exactness rests only on the statement in Section 5 that 'in the large N limit, the central limit theorem guarantees the Gaussianity of the process'; this is an uncontrolled asymptotic approximation and a legitimate correctness risk, but it is not a circular reduction — the effective equations do not presuppose the value of L_test. The only self-citation, Helias & Dahmen (2020), is used for the standard MSRDJ formalism and the disorder-averaging trick; the action is re-derived in Appendix A, so the citation is not load-bearing and does not force any result. No 'uniqueness theorem' from the authors' prior work is invoked to forbid alternatives, and no known empirical pattern is renamed as a prediction. The absence of error bars in the finite-N simulations (N=P=100) affects how strongly the quantitative match can be asserted, but it does not indicate that the theory's output was constructed from its inputs.
Assumptions & free parameters
assumptions (4)
- domain assumption Gaussian feature assumption: ψ_i(x) are i.i.d. Gaussian across samples with covariance η_i δ_ij, making Q a Wishart matrix.
- domain assumption Power-law spectral decay η_i = η_1 i^{-1-γ}.
- domain assumption Self-averaging: the data-averaged generating functional ⟨Z⟩ describes the typical single-realization dynamics.
- ad hoc to paper The variational Gaussian approximation becomes exact as N → ∞ because the central limit theorem guarantees Gaussianity of the effective process.
Cite this review
Pith. "Pith review of Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues." pith.science (2026). https://pith.science/paper/V3OVCADL
@misc{pith2026260223039,
author = {Pith},
title = {Pith review of: Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues},
year = {2026},
howpublished = {\url{https://pith.science/paper/V3OVCADL}},
note = {Machine review of arXiv:2602.23039}
}
read the original abstract
Training large neural networks exposes neural scaling laws for the generalization error, which points to a universal behavior across network architectures of learning in high dimensions. It was also shown that this effect persists in the limit of highly overparametrized networks as well as the Neural network Gaussian process limit. We here develop a principled understanding of the typical behavior of generalization in Neural Network Gaussian process regression dynamics. We derive a dynamical mean-field theory that captures the typical case learning dynamics: This allows us to unify multiple existing regimes of learning studied in the current literature, namely Bayesian inference on Gaussian processes, gradient flow with or without weight-decay, and stochastic Langevin training dynamics. Employing tools from statistical physics, the unified framework we derive in either of these cases yields an effective description of the high-dimensional microscopic behavior of networks dynamics in terms of lower dimensional order parameters. We show that collective training dynamics may be separated into the dynamics of N independent eigenmodes, those evolution equations are only coupled through collective response functions and a common statistics of an effective, independent noise. Our approach allows us to quantitatively explain the dynamics of the generalization error by linking spectral and dynamical properties of learning on data with power law spectra, including phenomena such as neural scaling laws and the effect of early stopping.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
Asymmetric Scaling Laws from Sparse Features
A sparse-activation model predicts double-descent loss with distinct under- and over-parameterized scaling exponents set by sparsity, plus a compute-optimal frontier favoring dataset growth.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Advani, M. S., Saxe, A. M., and Sompolinsky, H. High-dimensional dynamics of generalization error in neural networks. Neural Networks, 132: 0 428--446, 2020. doi:10.1016/j.neunet.2020.08.022
-
[3]
A dynamical model of neural scaling laws
Bordelon, B., Atanasov, A., and Pehlevan, C. A dynamical model of neural scaling laws. In Proceedings of the 41st International Conference on Machine Learning , volume 235 of ICML '24 , pp.\ 4345--4382. JMLR.org, 2024
2024
-
[4]
How feature learning can improve neural scaling laws
Bordelon, B., Atanasov, A., and Pehlevan, C. How feature learning can improve neural scaling laws. Journal of Statistical Mechanics Theory and Experiment, 2025. doi:10.1088/1742-5468/adefb1
-
[5]
Bulanadi, R. and Paruch, P. Identifying and analyzing power-law scaling in two-dimensional image datasets. Physical Review E, 109 0 (6): 0 064135, 2024. doi:10.1103/PhysRevE.109.064135
-
[6]
Canatar, A., Bordelon, B., and Pehlevan, C. Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks. Nature Communications, 2021. doi:10.1038/s41467-021-23103-1
-
[7]
Coppola, G. P., Helias, M., and Ringel, Z. Renormalization group for deep neural networks: Universality of learning and scaling laws, 2026. URL https://arxiv.org/abs/2510.25553
arXiv 2026
-
[8]
Techniques de renormalisation de la thÉorie des champs et dynamique des phÉnomÈnes critiques
De Dominicis, C. Techniques de renormalisation de la thÉorie des champs et dynamique des phÉnomÈnes critiques. Journal de Physique Colloques, 37 0 (C1): 0 C1, 1976. doi:10.1051/jphyscol:1976138
Show all 36 references
-
[9]
Dynamics as a substitute for replicas in systems with quenched random impurities
De Dominicis, C. Dynamics as a substitute for replicas in systems with quenched random impurities. Physical Review B, 18 0 (9): 0 4913, 1978
1978
-
[10]
Dunmur, A. P. and Wallace, D. J. Learning and generalization in a linear perceptron stochastically trained with noisy data. Journal of Physics A: Mathematical and General, 26 0 (21): 0 5767, nov 1993. doi:10.1088/0305-4470/26/21/016. URL https://doi.org/10.1088/0305-4470/26/21/016
1993 doi
-
[11]
Halkj r, S. r. and Winther, O. The effect of correlated input data on the dynamics of learning. In Mozer, M., Jordan, M., and Petsche, T. (eds.), Advances in Neural Information Processing Systems, volume 9. MIT Press, 1996. URL https://proceedings.neurips.cc/paper_files/paper/...
1996
-
[12]
and Dahmen, D
Helias, M. and Dahmen, D. Statistical Field Theory for Neural Networks. Springer International Publishing, 2020. doi:10.1007/978-3-030-46444-8
2020 doi
-
[13]
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems 31, pp.\ 8580--8589, Long Beach, CA, USA, 2018. URL https://proceedings.neurips.cc/paper/2018/file/5a4be1fa34e...
2018
-
[14]
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. arxiv, pp.\ 1806.07572, 2020
2020 arXiv
-
[15]
On a lagrangean for classical field dynamics and renormalization group calculations of dynamical critical properties
Janssen, H.-K. On a lagrangean for classical field dynamics and renormalization group calculations of dynamical critical properties. Zeitschrift f\"ur Physik B Condensed Matter and Quanta, 23 0 (4): 0 377--380, 1976. doi:10.1007/BF01316547
1976 doi
-
[16]
B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[17]
1/f2 Characteristics and Isotropy in the Fourier Power Spectra of Visual Art , Cartoons , Comics , Mangas , and Different Categories of Photographs
Koch, M., Denzler, J., and Redies, C. 1/f2 Characteristics and Isotropy in the Fourier Power Spectra of Visual Art , Cartoons , Comics , Mangas , and Different Categories of Photographs . PLOS ONE, 5 0 (8): 0 e12268, 2010. doi:10.1371/journal.pone.0012268
2010 doi
-
[18]
Learning with noise in a linear perceptron
Krogh, A. Learning with noise in a linear perceptron. Journal of Physics A: Mathematical and General, 1992. doi:10.1088/0305-4470/25/5/019
1992 doi
-
[19]
and Hertz, J
Krogh, A. and Hertz, J. Dynamics of generalization in linear perceptrons. In Lippmann, R., Moody, J., and Touretzky, D. (eds.), Advances in Neural Information Processing Systems, volume 3. Morgan-Kaufmann, 1990. URL https://proceedings.neurips.cc/paper_files/paper/1990/file/0b...
1990
-
[20]
and Hertz, J
Krogh, A. and Hertz, J. Generalization in a linear perceptron in the presence of noise. Journal of Physics A: Mathematical and General, 1992. doi:10.1088/0305-4470/25/5/020
1992 doi
- [21]
-
[22]
Deep neural networks as gaussian processes
Lee, J., Sohl-Dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y. Deep neural networks as gaussian processes. In International Conference on Learning Representations, Vancouver, British Columbia, Canada, 2018. OpenReview.net. URL https://openreview.net/foru...
2018
-
[23]
and Sompolinsky, H
Li, Q. and Sompolinsky, H. Statistical Mechanics of Deep Linear Neural Networks: The Backpropagating Kernel Renormalization . Physical Review X, 11 0 (3): 0 031059, 2021. doi:10.1103/PhysRevX.11.031059. URL https://journals.aps.org/prx/abstract/10.1103/PhysRevX.11.031059
2021 doi
- [24]
-
[25]
Statistical dynamics of classical systems
Martin, P., Siggia, E., and Rose, H. Statistical dynamics of classical systems. Physical Review A, 8 0 (1): 0 423--437, 1973
1973
-
[26]
Deep double descent: where bigger models and more data hurt*
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I. Deep double descent: where bigger models and more data hurt*. Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12): 0 124003, 2021. doi:10.1088/1742-5468/ac3a74. URL https://doi.org/10...
2021 doi
-
[27]
Predicting the outputs of finite networks trained with journal = arXiv ,noisy gradients
Naveh, G., Ben-David, O., Sompolinsky, H., and Ringel, Z. Predicting the outputs of finite networks trained with journal = arXiv ,noisy gradients . arxiv, 2020
2020
-
[28]
Neal, R. M. Bayesian learning for neural networks. Springer, 1995
1995
-
[29]
A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit
Pacelli, R., Ariosto, S., Pastore, M., Ginelli, F., Gherardi, M., and Rotondo, P. A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit. Nature Machine Intelligence, 5 0 (12): 0 1497--1507, December 2023. ISSN 2522-5839. doi:10.103...
2023 doi
-
[30]
and Recht, B
Rahimi, A. and Recht, B. Random features for large-scale kernel machines. In Platt, J., Koller, D., Singer, Y., and Roweis, S. (eds.), Advances in Neural Information Processing Systems, volume 20. Curran Associates, Inc., 2007. URL https://proceedings.neurips.cc/paper_files/pa...
2007
-
[31]
From kernels to features: A multi-scale adaptive theory of feature learning
Rubin, N., Fischer, K., Lindner, J., Dahmen, D., Seroussi, I., Ringel, Z., Krämer, M., and Helias, M. From kernels to features: A multi-scale adaptive theory of feature learning. arXiv preprint arXiv:2502.03210, 2025
2025 arXiv
-
[32]
Schaaf, A. v. d. and Hateren, J. H. v. Modelling the Power Spectra of Natural Images : Statistics and Information . Vision Research, 36 0 (17): 0 2759--2770, 1996. doi:10.1016/0042-6989(96)00002-8
1996 doi
-
[33]
Separation of scales and a thermodynamic description of feature learning in some cnns
Seroussi, I., Naveh, G., and Ringel, Z. Separation of scales and a thermodynamic description of feature learning in some cnns. Nature Communications, 14 0 (1): 0 908, February 2023. ISSN 2041-1723. doi:10.1038/s41467-023-36361-y
2023 doi
-
[34]
S., Sompolinsky, H., and Tishby, N
Seung, H. S., Sompolinsky, H., and Tishby, N. Statistical mechanics of learning from examples. Physical Review A, 45 0 (8): 0 6056--6091, 1992. doi:10.1103/PhysRevA.45.6056
1992 doi
-
[35]
Learning in large linear perceptrons and why the thermodynamic limit is relevant to the real world
Sollich, P. Learning in large linear perceptrons and why the thermodynamic limit is relevant to the real world. In Tesauro, G., Touretzky, D., and Leen, T. (eds.), Advances in Neural Information Processing Systems, volume 7. MIT Press, 1994. URL https://proceedings.neurips.cc/...
1994
-
[36]
Computing with infinite networks
Williams, C. Computing with infinite networks. Neural Information Processing Systems, 1996
1996
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.