REVIEW 12 cited by
A Solvable Model of Neural Scaling Laws
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically, their performance behaves predictably as a power law in either parameters or dataset size until bottlenecked by the other resource. To understand this better, we first identify the necessary properties allowing such scaling laws to arise and then propose a statistical model -- a joint generative data model and random feature model -- that captures this neural scaling phenomenology. By solving this model in the dual limit of large training set size and large number of parameters, we gain insight into (i) the statistical structure of datasets and tasks that lead to scaling laws, (ii) the way nonlinear feature maps, such as those provided by neural networks, enable scaling laws when trained on these datasets, (iii) the optimality of the equiparameterization scaling of training sets and parameters, and (iv) whether such scaling laws can break down and how they behave when they do. Key findings are the manner in which the power laws that occur in the statistics of natural datasets are extended by nonlinear random feature maps and then translated into power-law scalings of the test loss and how the finite extent of the data's spectral power law causes the model's performance to plateau.
Forward citations
Cited by 12 Pith papers
-
Smooth Scaling Laws Hide Stepwise Token Learning
Token loss trajectories follow localized sigmoids whose learning-time spectrum quantitatively reconstructs scaling-law derivatives on T, D, and M axes and enables faster training via distribution reshaping.
-
Explaining Data Mixing Scaling Laws
Under a shared-head/disjoint-tail assumption, multi-domain loss decomposes into a capacity-competition term c_i x_i^*(h)^{-b_i} plus a per-domain noise term A_i(Dh_i)^{-a_i}, and the fitted law extrapolates optimal mi...
-
Universal One-third Time Scaling in Learning Peaked Distributions
Softmax + cross-entropy on peaked targets yields loss ∼ t^{−1/3}, giving an architecture-driven explanation for power-law LLM training time without power-law data.
-
Dimension-adapted Momentum Outscales SGD
DANA, with dimension- and time-dependent momentum, provably outscales SGD on power-law random features when 2α>1, improving loss exponents and compute-optimal curves.
-
Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail
Neural scaling occurs because larger models maintain learning on weaker eigenmodes of the eNTK that smaller models cannot access.
-
Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.
-
Inverse Depth Scaling From Most Layers Being Similar
LLM loss decreases roughly inversely with depth because most layers act as a redundant ensemble that averages errors, not as a compositional hierarchy.
-
From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
Zipf's law, via differential Heaps and Hilberg laws, forces a power-law lower bound on the excess cross entropy of any entropy-bounded foundation model.
-
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
For diagonal and quadratic two-layer networks, training maps to LASSO and matrix compressed sensing, yielding a full phase diagram of excess-risk scaling exponents and a spectral characterization of the trained weights.
-
Models of Heavy-Tailed Mechanistic Universality
A new random matrix model with one structure parameter explains heavy-tailed spectra in trained networks, and yields scaling laws, optimizer-tail behavior, and a description of the five-plus-one phases of training.
-
X-Factor: Quality Is a Dataset-Intrinsic Property
Across 2,500 class-balanced MNIST subsets and 10 model architectures, test-error Z-scores correlate strongly across models (mean R2=0.82 excluding GNB), supporting dataset quality as an intrinsic property.
-
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency
The paper's finding that training efficiency declines monotonically with token count is guaranteed by its efficiency metric, which divides by token count and power consumption.
Discussion (0). Sign in to comment.