Pith. sign in

REVIEW 12 cited by

A Solvable Model of Neural Scaling Laws

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.16859 v1 pith:CSO363QM submitted 2022-10-30 cs.LG hep-thstat.ML

classification cs.LGhep-thstat.ML
keywords scalinglawsmodelneuralparametersdatasetsfeaturelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically, their performance behaves predictably as a power law in either parameters or dataset size until bottlenecked by the other resource. To understand this better, we first identify the necessary properties allowing such scaling laws to arise and then propose a statistical model -- a joint generative data model and random feature model -- that captures this neural scaling phenomenology. By solving this model in the dual limit of large training set size and large number of parameters, we gain insight into (i) the statistical structure of datasets and tasks that lead to scaling laws, (ii) the way nonlinear feature maps, such as those provided by neural networks, enable scaling laws when trained on these datasets, (iii) the optimality of the equiparameterization scaling of training sets and parameters, and (iv) whether such scaling laws can break down and how they behave when they do. Key findings are the manner in which the power laws that occur in the statistics of natural datasets are extended by nonlinear random feature maps and then translated into power-law scalings of the test loss and how the finite extent of the data's spectral power law causes the model's performance to plateau.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Smooth Scaling Laws Hide Stepwise Token Learning

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    Token loss trajectories follow localized sigmoids whose learning-time spectrum quantitatively reconstructs scaling-law derivatives on T, D, and M axes and enables faster training via distribution reshaping.

  2. Explaining Data Mixing Scaling Laws

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Under a shared-head/disjoint-tail assumption, multi-domain loss decomposes into a capacity-competition term c_i x_i^*(h)^{-b_i} plus a per-domain noise term A_i(Dh_i)^{-a_i}, and the fitted law extrapolates optimal mi...

  3. Universal One-third Time Scaling in Learning Peaked Distributions

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Softmax + cross-entropy on peaked targets yields loss ∼ t^{−1/3}, giving an architecture-driven explanation for power-law LLM training time without power-law data.

  4. Dimension-adapted Momentum Outscales SGD

    stat.ML 2025-05 conditional novelty 7.0 of 10

    DANA, with dimension- and time-dependent momentum, provably outscales SGD on power-law random features when 2α>1, improving loss exponents and compute-optimal curves.

  5. Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Neural scaling occurs because larger models maintain learning on weaker eigenmodes of the eNTK that smaller models cannot access.

  6. Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues

    cond-mat.dis-nn 2026-02 conditional novelty 6.0 of 10

    A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.

  7. Inverse Depth Scaling From Most Layers Being Similar

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LLM loss decreases roughly inversely with depth because most layers act as a redundant ensemble that averages errors, not as a compositional hierarchy.

  8. From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis

    cs.IT 2025-12 conditional novelty 6.0 of 10

    Zipf's law, via differential Heaps and Hilberg laws, forces a power-law lower bound on the excess cross entropy of any entropy-bounded foundation model.

  9. Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

    cs.LG 2025-09 conditional novelty 6.0 of 10

    For diagonal and quadratic two-layer networks, training maps to LASSO and matrix compressed sensing, yielding a full phase diagram of excess-risk scaling exponents and a spectral characterization of the trained weights.

  10. Models of Heavy-Tailed Mechanistic Universality

    stat.ML 2025-06 conditional novelty 6.0 of 10

    A new random matrix model with one structure parameter explains heavy-tailed spectra in trained networks, and yields scaling laws, optimizer-tail behavior, and a description of the five-plus-one phases of training.

  11. X-Factor: Quality Is a Dataset-Intrinsic Property

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Across 2,500 class-balanced MNIST subsets and 10 model architectures, test-error Z-scores correlate strongly across models (mean R2=0.82 excluding GNB), supporting dataset quality as an intrinsic property.

  12. Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

    cs.LG 2026-01 reject novelty 4.0 of 10

    The paper's finding that training efficiency declines monotonically with token count is guaranteed by its efficiency metric, which divides by token count and power consumption.

Pith tools