Pith. sign in

REVIEW 18 cited by

Deep Double Descent: Where Bigger Models and More Data Hurt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02292 v1 pith:NAFXKFON submitted 2019-12-04 cs.LG cs.CVcs.NEstat.ML

classification cs.LGcs.CVcs.NEstat.ML
keywords modelcomplexitydescentdoubledeepfunctiongetsmeasure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that a variety of modern deep learning tasks exhibit a "double-descent" phenomenon where, as we increase model size, performance first gets worse and then gets better. Moreover, we show that double descent occurs not just as a function of model size, but also as a function of the number of training epochs. We unify the above phenomena by defining a new complexity measure we call the effective model complexity and conjecture a generalized double descent with respect to this measure. Furthermore, our notion of model complexity allows us to identify certain regimes where increasing (even quadrupling) the number of train samples actually hurts test performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

    cs.LG 2022-01 unverdicted novelty 8.0 of 10

    Neural networks exhibit grokking on small algorithmic datasets, achieving perfect generalization well after overfitting.

  2. How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    The authors derive a Maximally Scale-Stable Parameterization (MSSP) for MoE models that achieves robust learning-rate transfer and monotonic performance gains with scale across co-scaling regimes of width, experts, an...

  3. Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A fast hash-based simplification engine enables a transformer-based symbolic regression system to train on 512M simplified expressions and match PySR on the FastSRB benchmark with better parsimony scaling.

  4. On Spectral Properties of Gradient-based Explanation Methods

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.

  5. Detecting AI Assistance in Abstract Complex Tasks

    cs.AI 2025-07 reject novelty 6.0 of 10

    Converting behavioral search traces into image channels plus an exploration/exploitation time series lets a small CNN-RNN detect AI assistance with about 86% accuracy on a balanced lab task.

  6. Language Models (Mostly) Know What They Know

    cs.CL 2022-07 unverdicted novelty 6.0 of 10

    Language models show good calibration when asked to estimate the probability that their own answers are correct, with performance improving as models get larger.

  7. Scaling Laws and Interpretability of Learning from Repeated Data

    cs.LG 2022-05 accept novelty 6.0 of 10

    Repeating 0.1% of training data 100 times degrades an 800M parameter model's performance to that of a 400M model by damaging copying mechanisms and induction heads associated with generalization.

  8. A General Language Assistant as a Laboratory for Alignment

    cs.CL 2021-12 conditional novelty 6.0 of 10

    Ranked preference modeling outperforms imitation learning for language model alignment and scales more favorably with model size.

  9. Scaling Laws for Transfer

    cs.LG 2021-02 unverdicted novelty 6.0 of 10

    Effective data transferred from pre-training to fine-tuning is described by a power law in model parameter count and fine-tuning dataset size, acting like a multiplier on the fine-tuning data.

  10. Asymmetric Scaling Laws from Sparse Features

    stat.ML 2026-05 unverdicted novelty 5.0 of 10

    A sparse-activation model predicts double-descent loss with distinct under- and over-parameterized scaling exponents set by sparsity, plus a compute-optimal frontier favoring dataset growth.

  11. Double Descent and Overparameterization in Particle Physics Data

    hep-ex 2025-09 conditional novelty 5.0 of 10

    Double descent, a test-error peak followed by recovery at high capacity, appears in jet regression and event classification on ATLAS open data; with early stopping, overparameterized models can beat classical ones.

  12. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  13. A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    Repeated measures experiment on small LM training shows validation loss drops then rises under fixed token budget with significant interval effects and recurrent backslides.

  14. Unified Neural Scaling Laws

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    Presents a single functional form for neural scaling that unifies multiple scaling dimensions and claims higher extrapolation accuracy than prior forms across diverse tasks and architectures.

  15. Position: Ideas Should be the Center of Machine Learning Research

    cs.LG 2026-05 conditional novelty 4.0 of 10

    Machine learning research should prioritize ideas by testing their predicted behavioral signatures in modern models through custom experiments instead of leaderboard chasing or abstract theorems.

  16. Does Order Matter : Connecting The Law of Robustness to Robust Generalization

    cs.LG 2026-02 reject novelty 4.0 of 10

    The paper proves R(ℓρ∘B_L∘S) ≤ 8R(B_L∘S) but does not derive the advertised Ω(n^{1/d}) recovery or the missing local-scale result.

  17. Optimizers Qualitatively Alter Solutions And We Should Leverage This

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Deep learning optimizers should be designed to induce desired solution properties, not just convergence speed; different optimizers demonstrably land in qualitatively different minima.

  18. Six Open Questions in Machine-Learned Interatomic Potential Foundation Models

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 2.0 of 10

    This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.

Pith tools