Pith. sign in

REVIEW 13 cited by

Deep Double Descent: Where Bigger Models and More Data Hurt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02292 v1 pith:NAFXKFON submitted 2019-12-04 cs.LG cs.CVcs.NEstat.ML

Deep Double Descent: Where Bigger Models and More Data Hurt

classification cs.LG cs.CVcs.NEstat.ML
keywords modelcomplexitydescentdoubledeepfunctiongetsmeasure
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We show that a variety of modern deep learning tasks exhibit a "double-descent" phenomenon where, as we increase model size, performance first gets worse and then gets better. Moreover, we show that double descent occurs not just as a function of model size, but also as a function of the number of training epochs. We unify the above phenomena by defining a new complexity measure we call the effective model complexity and conjecture a generalized double descent with respect to this measure. Furthermore, our notion of model complexity allows us to identify certain regimes where increasing (even quadrupling) the number of train samples actually hurts test performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

    cs.LG 2022-01 unverdicted novelty 8.0

    Neural networks exhibit grokking on small algorithmic datasets, achieving perfect generalization well after overfitting.

  2. How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

    cs.LG 2026-05 unverdicted novelty 7.0

    The authors derive a Maximally Scale-Stable Parameterization (MSSP) for MoE models that achieves robust learning-rate transfer and monotonic performance gains with scale across co-scaling regimes of width, experts, an...

  3. Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression

    cs.LG 2026-02 conditional novelty 6.0

    A fast hash-based simplification engine enables a transformer-based symbolic regression system to train on 512M simplified expressions and match PySR on the FastSRB benchmark with better parsimony scaling.

  4. Language Models (Mostly) Know What They Know

    cs.CL 2022-07 unverdicted novelty 6.0

    Language models show good calibration when asked to estimate the probability that their own answers are correct, with performance improving as models get larger.

  5. Scaling Laws and Interpretability of Learning from Repeated Data

    cs.LG 2022-05 accept novelty 6.0

    Repeating 0.1% of training data 100 times degrades an 800M parameter model's performance to that of a 400M model by damaging copying mechanisms and induction heads associated with generalization.

  6. A General Language Assistant as a Laboratory for Alignment

    cs.CL 2021-12 conditional novelty 6.0

    Ranked preference modeling outperforms imitation learning for language model alignment and scales more favorably with model size.

  7. Scaling Laws for Transfer

    cs.LG 2021-02 unverdicted novelty 6.0

    Effective data transferred from pre-training to fine-tuning is described by a power law in model parameter count and fine-tuning dataset size, acting like a multiplier on the fine-tuning data.

  8. Asymmetric Scaling Laws from Sparse Features

    stat.ML 2026-05 unverdicted novelty 5.0

    A sparse-activation model predicts double-descent loss with distinct under- and over-parameterized scaling exponents set by sparsity, plus a compute-optimal frontier favoring dataset growth.

  9. A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget

    cs.AI 2026-06 unverdicted novelty 4.0

    Repeated measures experiment on small LM training shows validation loss drops then rises under fixed token budget with significant interval effects and recurrent backslides.

  10. Unified Neural Scaling Laws

    cs.LG 2026-05 unverdicted novelty 4.0

    Presents a single functional form for neural scaling that unifies multiple scaling dimensions and claims higher extrapolation accuracy than prior forms across diverse tasks and architectures.

  11. Position: Ideas Should be the Center of Machine Learning Research

    cs.LG 2026-05 conditional novelty 4.0

    Machine learning research should prioritize ideas by testing their predicted behavioral signatures in modern models through custom experiments instead of leaderboard chasing or abstract theorems.

  12. Does Order Matter : Connecting The Law of Robustness to Robust Generalization

    cs.LG 2026-02 reject novelty 4.0

    The paper proves R(ℓρ∘B_L∘S) ≤ 8R(B_L∘S) but does not derive the advertised Ω(n^{1/d}) recovery or the missing local-scale result.

  13. Six Open Questions in Machine-Learned Interatomic Potential Foundation Models

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 2.0

    This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.