Pith. sign in

REVIEW 9 cited by

Learning Curve Theory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.04074 v1 pith:NZLRBBGZ submitted 2021-02-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords datascalingbetadecreaseserrorlawslearningpower
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recently a number of empirical "universal" scaling law papers have been published, most notably by OpenAI. `Scaling laws' refers to power-law decreases of training or test error w.r.t. more data, larger neural networks, and/or more compute. In this work we focus on scaling w.r.t. data size $n$. Theoretical understanding of this phenomenon is largely lacking, except in finite-dimensional models for which error typically decreases with $n^{-1/2}$ or $n^{-1}$, where $n$ is the sample size. We develop and theoretically analyse the simplest possible (toy) model that can exhibit $n^{-\beta}$ learning curves for arbitrary power $\beta>0$, and determine whether power laws are universal or depend on the data distribution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Smooth Scaling Laws Hide Stepwise Token Learning

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    Token loss trajectories follow localized sigmoids whose learning-time spectrum quantitatively reconstructs scaling-law derivatives on T, D, and M axes and enables faster training via distribution reshaping.

  2. Universal One-third Time Scaling in Learning Peaked Distributions

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Softmax + cross-entropy on peaked targets yields loss ∼ t^{−1/3}, giving an architecture-driven explanation for power-law LLM training time without power-law data.

  3. Inverse Depth Scaling From Most Layers Being Similar

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LLM loss decreases roughly inversely with depth because most layers act as a redundant ensemble that averages errors, not as a compositional hierarchy.

  4. Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Loss deceleration, a piecewise-linear break in log-log loss curves, is attributed to zero-sum learning where per-example gradients oppose one another, and scaling helps by mitigating it.

  5. Scaling Pre-training to One Hundred Billion Data for Vision Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Scaling VLM pretraining from 10B to 100B image-text pairs yields saturation on standard benchmarks but large gains on cultural diversity, low-resource language retrieval, and subgroup disparity.

  6. Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems

    cs.AI 2025-02 conditional novelty 6.0 of 10

    Recursively applying the first half of a transformer before the second half (RINS) improves language modeling and vision-language accuracy under compute-matched comparisons.

  7. Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Classification accuracy on inertial HAR and SLR tasks follows a consistent logarithmic growth with training-set size, enabling a MAPD-based stability-point metric that often saturates far below traditional heuristics.

  8. Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development

    cs.LG 2026-07 conditional novelty 5.0 of 10

    In a stylized model, a proactive flywheel that fixes whole groups of related scenarios needs Θ(K log K) update rounds versus Θ(M log M) for reactive patching.

  9. Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Using NTK trace and effective rank, this paper shows that model and data scaling improve test loss at similar rates but drive internal dynamics in opposite directions, and estimates a feature-learning width limit well...

Pith tools