Pith. sign in

REVIEW 14 cited by

A Constructive Prediction of the Generalization Error Across Scales

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.12673 v2 pith:G4WMOLK4 submitted 2019-09-27 cs.LG cs.CLcs.CVstat.ML

A Constructive Prediction of the Generalization Error Across Scales

classification cs.LG cs.CLcs.CVstat.ML
keywords modelformscalesacrossdataerrorgeneralizationdependency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The dependency of the generalization error of neural networks on model and dataset size is of critical importance both in practice and for understanding the theory of neural networks. Nevertheless, the functional form of this dependency remains elusive. In this work, we present a functional form which approximates well the generalization error in practice. Capitalizing on the successful concept of model scaling (e.g., width, depth), we are able to simultaneously construct such a form and specify the exact models which can attain it across model/data scales. Our construction follows insights obtained from observations conducted over a range of model/data scales, in various model types and datasets, in vision and language tasks. We show that the form both fits the observations well across scales, and provides accurate predictions from small- to large-scale models and data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Scaling Laws for Neural Language Models

    cs.LG 2020-01 unverdicted novelty 8.0

    Empirical power-law scaling governs language model loss versus model size, data size, and compute, enabling optimal allocation of training compute.

  2. Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis

    cs.LG 2026-05 unverdicted novelty 7.0

    QuBD extends algorithmic complexity estimation to quantized DNN weights, revealing that complexity decreases during learning, increases with overfitting, follows grokking patterns, and correlates with generalization.

  3. Scaling Laws for Autoregressive Generative Modeling

    cs.LG 2020-10 accept novelty 7.0

    Autoregressive transformers follow power-law scaling laws for cross-entropy loss with nearly universal exponents relating optimal model size to compute budget across four domains.

  4. Label-Efficient Dataset Pruning via Semi-Supervised Pseudo-Labeling

    cs.LG 2026-05 unverdicted novelty 6.0

    SemiPrune uses a small labeled subset and semi-supervised pseudo-labeling to enable supervised dataset pruning methods, achieving state-of-the-art results on domain-specific, image-corrupted, and long-tailed datasets.

  5. A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks

    cs.LG 2026-05 unverdicted novelty 6.0

    Depth expansion in normalized residual networks yields provable test-risk improvement through representational, optimization, and generalization gains under first-order descent and norm-control conditions.

  6. Language Models (Mostly) Know What They Know

    cs.CL 2022-07 unverdicted novelty 6.0

    Language models show good calibration when asked to estimate the probability that their own answers are correct, with performance improving as models get larger.

  7. A General Language Assistant as a Laboratory for Alignment

    cs.CL 2021-12 conditional novelty 6.0

    Ranked preference modeling outperforms imitation learning for language model alignment and scales more favorably with model size.

  8. Scaling Laws for Transfer

    cs.LG 2021-02 unverdicted novelty 6.0

    Effective data transferred from pre-training to fine-tuning is described by a power law in model parameter count and fine-tuning dataset size, acting like a multiplier on the fine-tuning data.

  9. Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification

    cs.LG 2026-07 conditional novelty 5.0

    Classification accuracy on inertial HAR and SLR tasks follows a consistent logarithmic growth with training-set size, enabling a MAPD-based stability-point metric that often saturates far below traditional heuristics.

  10. Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization

    cs.LG 2026-05 unverdicted novelty 5.0

    Tuning the depth-width ratio positions models in an efficient neural interaction interval that correlates with better generalization under fixed budgets and remains stable with scale.

  11. The Platonic Representation Hypothesis

    cs.LG 2024-05 unverdicted novelty 5.0

    Representations learned by large AI models are converging toward a shared statistical model of reality.

  12. Compute-Optimal Network Design for Echocardiography Myocardial Segmentation and Perfusion Quantification using Neural Scaling Laws

    eess.IV 2026-06 unverdicted novelty 4.0

    Neural scaling laws fitted to subset performance on CAMUS and CEUS echocardiography datasets enable selection of smaller networks achieving state-of-the-art myocardial segmentation with 240-fold parameter reduction an...

  13. Unified Neural Scaling Laws

    cs.LG 2026-05 unverdicted novelty 4.0

    Presents a single functional form for neural scaling that unifies multiple scaling dimensions and claims higher extrapolation accuracy than prior forms across diverse tasks and architectures.

  14. Unifying Learning Dynamics and Generalization in Transformers Scaling Law

    cs.LG 2025-12 reject novelty 4.0

    Claims a two-stage transformer scaling law (exponential then C^{-1/6}) with matching bounds, but the lower bounds are missing, the exponent is inconsistent (-1/7 vs -1/6), and the law is an artifact of hand-set M = Θ(...