Pith. sign in

REVIEW 2 major objections 1 minor 117 references

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read The inductive bias of a trained neural network is its architectural prior filtered by the forgetting dynamics of the training pipeline.

desk verdict Low-LR SGD keeps a 26-point test-acc spread across init scales on CIFAR ResNets even at 99.5% train acc, while Adam and L2 erase it, but the spread may trace to other trajectory differences rather than retained initialization memory. read the letter →

arxiv 2605.29152 v1 pith:MLZBYQUY submitted 2026-05-27 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords initializationmemoryinductivebiasneuralnetworktrainingSGDforgettingregularizationResNetCIFAR-10
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Neural networks start with a prior induced by random initialization, but training changes how much of that survives. Experiments on ResNets for CIFAR-10 reveal that low-learning-rate SGD keeps strong dependence on initialization scale, producing up to 26.5 percentage point swings in test accuracy even when training accuracy exceeds 99.5 percent. Adam optimizers and explicit regularization largely remove this dependence by accelerating forgetting of the initial conditions. The result is that effective inductive bias is not architecture alone but architecture after the training process has filtered the starting point.

What carries the argument

Initialization memory, the dependence of the validation-selected predictor on the scale of the random initialization; it quantifies how much initial bias survives training.

What would settle it

An experiment in which the accuracy spread across init scales disappears when other optimization factors are controlled while maintaining the same training accuracy.

Watch

Extended reading notes

Core claim

In controlled experiments, initialization memory—the dependence of the final predictor on random initialization scale—persists under gradient-flow-like dynamics such as low-LR SGD but is erased on timescales set by stochastic effects, norm decay, or adaptive preconditioning; therefore the practical bias equals the architectural prior after filtering by forgetting dynamics, and regularizers improve generalization precisely by erasing initialization memory.

Load-bearing premise

Variation in test accuracy across initialization scales after high training accuracy isolates retained initialization memory rather than other differences in optimization.

Editorial extensions

If this is right

  • Low-learning-rate SGD interpolates yet retains initialization memory, leading to large test accuracy variation.
  • Adam-family methods erase the dependence on initialization scale.
  • Pairing larger learning rates with L2 norm control causes SGD to forget initialization.
  • The time scale of forgetting is governed by the size of explicit or implicit regularization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Initialization scale may need to be tuned differently depending on the optimizer used.
  • This forgetting view could explain why certain training choices improve generalization beyond what architecture alone predicts.
  • Extending training time does not necessarily increase forgetting if the regime preserves memory.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper claims that trained deep networks retain a measurable dependence on the scale of their random initialization ('initialization memory'), which survives low-learning-rate SGD but is erased by Adam-family optimizers or explicit L2 regularization. This is demonstrated via controlled CIFAR-10 ResNet experiments showing a 26.5 percentage point spread in test accuracy across initialization scales at >=99.5% training accuracy; the spread persists even after extending training to 5000 epochs. The authors interpret the results as evidence that practical inductive bias is the architectural prior filtered by the forgetting dynamics of the training pipeline, with the same regularizers that aid generalization also erasing initialization memory.

Significance. If the reported accuracy spreads can be isolated to retained initialization memory rather than correlated differences in optimization trajectories, the work offers a useful empirical lens on how training dynamics shape effective priors. The controlled regime comparisons (low-LR SGD vs. Adam vs. L2) and the observation that extended epochs do not close the gap provide concrete, falsifiable distinctions between regimes. The absence of free parameters or fitted models in the core measurements is a strength of the empirical design.

major comments (2)
  1. [Abstract] Abstract and the CIFAR-10 ResNet-9 experiments: the central attribution of the 26.5 pp test-accuracy spread to retained 'initialization memory' (rather than other init-scale-dependent trajectory factors such as effective step-size distribution, basin geometry, or final weight norms) is load-bearing for the forgetting-time interpretation, yet the reported controls (high training accuracy plus extension to 5000 epochs) do not constrain these alternatives. High train accuracy plus fixed epoch count is insufficient to isolate the claimed mechanism.
  2. [Abstract] The manuscript reports quantitative separation under controlled setups but provides no details on the number of independent runs, error bars, or exact hyperparameter sweep ranges for the 26.5 pp spread. Without these, it is impossible to assess whether the observed variation is statistically robust or sensitive to minor implementation choices.
minor comments (1)
  1. Notation for 'initialization memory' is introduced without a formal definition or equation; a precise mathematical statement of the dependence being measured would improve clarity.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the careful reading and constructive feedback. We respond to each major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract and the CIFAR-10 ResNet-9 experiments: the central attribution of the 26.5 pp test-accuracy spread to retained 'initialization memory' (rather than other init-scale-dependent trajectory factors such as effective step-size distribution, basin geometry, or final weight norms) is load-bearing for the forgetting-time interpretation, yet the reported controls (high training accuracy plus extension to 5000 epochs) do not constrain these alternatives. High train accuracy plus fixed epoch count is insufficient to isolate the claimed mechanism.

    Authors: We define initialization memory operationally as the dependence of the validation-selected predictor on initialization scale under a fixed training procedure. The reported experiments show that this dependence survives low-LR SGD even after ≥99.5% training accuracy is reached and after training is extended to 5000 epochs, while the same dependence is erased under Adam or explicit L2 regularization. These regime contrasts are the primary evidence for the forgetting-time view. We agree that additional measurements (e.g., final weight norms across scales or basin geometry) would further constrain alternative explanations and will add a limitations paragraph discussing this point in the revision. revision: partial

  2. Referee: [Abstract] The manuscript reports quantitative separation under controlled setups but provides no details on the number of independent runs, error bars, or exact hyperparameter sweep ranges for the 26.5 pp spread. Without these, it is impossible to assess whether the observed variation is statistically robust or sensitive to minor implementation choices.

    Authors: We agree that these experimental details are necessary. The 26.5 pp figure is computed from multiple independent runs that vary only the initialization scale (different random seeds), and we will include the exact number of runs, standard-error bars, and the precise hyperparameter ranges used for all reported regimes in the revised manuscript and supplementary material. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity; purely empirical measurements with no self-referential derivations

full rationale

The paper reports controlled CIFAR-10 experiments on ResNets that measure test-accuracy spread across initialization scales under fixed training-accuracy thresholds. No equations, fitted parameters, or predictions are defined in terms of the target quantities. No self-citations are invoked as load-bearing uniqueness theorems or ansatzes. The central observations (e.g., 26.5 pp spread under low-LR SGD) are direct empirical outputs, not reductions of any input by construction. The work is therefore self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 1 assumptions · 1 invented entities

The central claim rests on empirical measurements of accuracy variation rather than theoretical derivations; the only addition is the newly introduced measurable quantity.

assumptions (1)
  • domain assumption Standard assumptions underlying SGD and Adam dynamics on non-convex loss surfaces
    Invoked when interpreting why low-LR SGD preserves memory while Adam erases it.
invented entities (1)
  • initialization memory
    purpose: Quantify the retained dependence of the validation-selected predictor on random initialization scale
    Newly defined to make the survival of initial bias measurable; no independent evidence outside the paper's experiments.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias." pith.science (2026). https://pith.science/paper/MLZBYQUY

@misc{pith2026260529152,
  author       = {Pith},
  title        = {Pith review of: Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MLZBYQUY}},
  note         = {Machine review of arXiv:2605.29152}
}
abstract

Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survives the training pipeline. To make the question measurable, we introduce initialization memory: the dependence of the validation-selected predictor on the scale of the random initialization. We perform controlled CIFAR-10 experiments on ResNets where initialization memory already sharply separates training regimes. Low-learning-rate SGD can interpolate while still remembering its initialization: on ResNet-9 with batch size $b=128$, test accuracy varies by $26.5$ percentage points across initialization scales despite $\ge99.5\%$ training accuracy. This is not undertraining: extending the same low-learning-rate regime to $5{,}000$ epochs leaves the spread essentially unchanged. In contrast, Adam-family methods largely erase the dependence. SGD can also be made to forget when larger learning rates are paired with explicit $L_2$ norm control. We interpret these findings in terms of the time scale of forgetting: gradient-flow-like dynamics can preserve initialization memory, whereas stochastic finite-step effects, explicit norm decay, and adaptive preconditioning erase it on scales governed by the size of explicit or implicit regularization. The practical inductive bias of a trained network is therefore not the architectural prior alone, but the architectural prior after being filtered by the forgetting dynamics of the training pipeline; and the same regularizers that improve generalization are precisely those that erase memory of initialization.

Figures

Figures reproduced from arXiv: 2605.29152 by the authors.

Figure 1
Figure 1. SGD remembers initialization; Adam-family methods forget. ResNet-9 under a shared low-learning-rate training procedure. Each curve shows the mean over n = 10 seeds; shaded bands indicate the 10th − 90th percentile range. SGD interpolates, but its generalization gap grows with σw. Adam, AdamW, and Muon show substantially weaker dependence on σw. The norm panels show radial memory: SGD retains sensitivity to the initi… view at source ↗
Figure 2
Figure 2. Large-batch fixed-epoch regimes forget initialization more slowly. ResNet-9 test accuracy at τbest, averaged over n = 10 seeds. Each panel corresponds to one batch size b ∈ {16, 32, 64, 128, 256}; within each panel, rows are optimizers and columns are initialization scales. Cells marked × did not reach 99.5% mean training accuracy at the best-validation-loss checkpoint τbest. SGD shows a strong left-to-right degrada… view at source ↗
Figure 3
Figure 3. Interpolation is not forgetting. (a,b) Interpolation epoch τinterp versus σw. SGD’s interpolation time grows sharply with σw, especially at large batch sizes; Adam’s is nearly flat. (c,d) Repair gap ∆repair = ValAccτbest − ValAccτinterp . At large batch (b≥128) Adam-family methods continue to gain validation accuracy after interpolation (τbest ≫τinterp, ∆repair >0). At small batch (b≤64) they reach τbest before inte… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: What helps SGD forget initialization? Test accuracy vs. σw for (a) b=16, and (b) b=128; inset shows the spread (pp) per configuration, sorted ascending. Long training alone leaves the curves nearly unchanged; larger learning rates and especially moderate LR with explic…
Figure 5
Figure 5. Figure 5: Depth makes poor forgetting dynamics more damaging; pooling is a partial confound. Test accuracy versus σw for ResNet-9, R9-AvgPool, ResNet-56, and ResNet-110. Switching pooling strategy reduces ResNet-9 performance, but the 9-layer R9 remains more robust than ResNet-5…
Figure 6
Figure 6. Figure 6: Initialization sensitivity decays on regularization timescales, not epoch count. (a– d) |β(t)| versus epoch for the SGD family (left) and adaptive methods (right) at b=16 (top) and b=128 (bottom). Vanilla SGD (red) and SGD+momentum (orange) stay flat; Adam/AdamW/Muon d…
Figure 7
Figure 7. Figure 7: Train accuracy vs. initialization scale σw (same layout as [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Test accuracy vs. initialization scale σw for ResNet-9 under BatchNorm + augmentation (blue), LayerNorm + augmentation (orange), and the BatchNorm baseline without augmentation (grey dashed). Lines: mean over seeds; shaded bands: 10th–90th percentile. 24 [PITH_FULL_IM…
Figure 9
Figure 9. Figure 9: Optimizer trajectories on the scalar loss (ab − 1)2 , initialization (a0, b0) = (1, 6). All four panels use the same two learning rates η ∈ {0.01, 0.04} (red and orange, respectively); only the optimizer changes. The black curve is the manifold of global minima ab = 1;…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

117 extracted references · 8 canonical work pages

  1. [1]

    Under- standing deep learning (still) requires rethinking generalization.Communications of the ACM, 64(3):107–115, 2021

    Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Under- standing deep learning (still) requires rethinking generalization.Communications of the ACM, 64(3):107–115, 2021

  2. [2]

    Strogatz.Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering

    Steven H. Strogatz.Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering. Westview Press, 2 edition, 2015

  3. [3]

    Hirsch, Stephen Smale, and Robert L

    Morris W. Hirsch, Stephen Smale, and Robert L. Devaney.Differential Equations, Dynamical Systems, and an Introduction to Chaos. Academic Press, 3 edition, 2013

  4. [4]

    On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks

    Hancheng Min, Salma Tarmoun, René Vidal, and Enrique Mallada. On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks. InProceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 7760–7768. PMLR, 2021

  5. [5]

    On the role of initialization on the implicit bias in deep linear networks, 2024

    Oria Gruber and Haim Avron. On the role of initialization on the implicit bias in deep linear networks, 2024

  6. [6]

    Camargo, and Ard A

    Guillermo Valle-Pérez, Chico Q. Camargo, and Ard A. Louis. Deep learning generalizes because the parameter–function map is biased towards simple functions. InInternational Conference on Learning Representations, 2019

  7. [7]

    Chris Mingard, Henry Rees, Guillermo Valle-Pérez, and Ard A. Louis. Deep neural networks have an inbuilt occam’s razor.Nature Communications, 16:220, 2025

  8. [8]

    Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026

    Thomas Fink. Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026

Show all 117 references
  1. [9]

    Understanding the difficulty of training deep feedforward neural networks

    Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. InProceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, volume 9 ofProceedings of Machine Learning Research, pages 249–256. P...

  2. [10]

    Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. InProceedings of the IEEE International Conference on Computer Vision, pages 1026–1034, 2015

  3. [11]

    Exponential expressivity in deep neural networks through transient chaos

    Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli. Exponential expressivity in deep neural networks through transient chaos. InAdvances in Neural Information Processing Systems, volume 29, 2016

  4. [12]

    Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein

    Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. Deep information propagation. InInternational Conference on Learning Representations, 2017

  5. [13]

    Schoenholz, and Surya Ganguli

    Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli. Resurrecting the sigmoid in deep learning through dynamical isometry: Theory and practice. InAdvances in Neural Information Processing Systems, volume 30, 2017

  6. [14]

    How to start training: The effect of initialization and architecture

    Boris Hanin and David Rolnick. How to start training: The effect of initialization and architecture. InAdvances in Neural Information Processing Systems, volume 31, 2018

  7. [15]

    Which neural net architectures give rise to exploding and vanishing gradients? InAdvances in Neural Information Processing Systems, volume 31, 2018

    Boris Hanin. Which neural net architectures give rise to exploding and vanishing gradients? InAdvances in Neural Information Processing Systems, volume 31, 2018. 10

  8. [16]

    Schoenholz, and Jeffrey Pennington

    Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S. Schoenholz, and Jeffrey Pennington. Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks. InProceedings of the 35th International Conference on Machine L...

  9. [17]

    Schoenholz

    Minmin Chen, Jeffrey Pennington, and Samuel S. Schoenholz. Dynamical isometry and a mean field theory of RNNs: Gating enables signal propagation in recurrent neural networks. InProceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Ma...

  10. [18]

    Greg Yang and Edward J. Hu. Tensor programs IV: Feature learning in infinite-width neural networks. InProceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 11727–11737. PMLR, 2021

  11. [19]

    Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao

    Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor programs V: Tuning large neural networks via zero-shot hyperparameter transfer. InAdvances in Neural Information Processing Syst...

  12. [20]

    Self-consistent dynamical field theory of kernel evo- lution in wide neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2023(11):114009, 2023

    Blake Bordelon and Cengiz Pehlevan. Self-consistent dynamical field theory of kernel evo- lution in wide neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2023(11):114009, 2023

  13. [21]

    Dynamics of finite width kernel and prediction fluctua- tions in mean field neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2024(10):104021, 2024

    Blake Bordelon and Cengiz Pehlevan. Dynamics of finite width kernel and prediction fluctua- tions in mean field neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2024(10):104021, 2024

  14. [22]

    Deep linear network training dynamics from random initialization: Data, width, depth, and hyperparameter transfer

    Blake Bordelon and Cengiz Pehlevan. Deep linear network training dynamics from random initialization: Data, width, depth, and hyperparameter transfer. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceedings of Machine Learning Research,...

  15. [23]

    Adaptive kernel predictors from feature-learning infinite limits of neural networks

    Clarissa Lauditi, Blake Bordelon, and Cengiz Pehlevan. Adaptive kernel predictors from feature-learning infinite limits of neural networks. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceedings of Machine Learning Research, pages 3261...

  16. [24]

    Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping.CoRR, abs/2002.06305, 2020

  17. [25]

    Zuidema, and Stella R

    Oskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra, Hailey Schoelkopf, Willem H. Zuidema, and Stella R. Biderman. PolyPythias: Stability and outliers across fifty language model pre-training runs. InInternational Conference on Learning Representations, 2025

  18. [26]

    Convergence and divergence of language models under different random seeds

    Finlay Fehlauer, Kyle Mahowald, and Tiago Pimentel. Convergence and divergence of language models under different random seeds. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32982–32991, Suzhou, China,

  19. [27]

    Association for Computational Linguistics

  20. [28]

    SeedPrints: Fingerprints can even tell which seed your large language model was trained from

    Yao Tong, Haonan Wang, Siquan Li, Kenji Kawaguchi, and Tianyang Hu. SeedPrints: Fingerprints can even tell which seed your large language model was trained from. In International Conference on Learning Representations, 2026. Poster

  21. [29]

    Transformers are born biased: Structural inductive biases at random initialization and their practical consequences, 2026

    Siquan Li, Yao Tong, Haonan Wang, and Tianyang Hu. Transformers are born biased: Structural inductive biases at random initialization and their practical consequences, 2026

  22. [30]

    Physics of language models: Part 3.1, knowledge storage and extraction

    Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 1067–1077. PMLR, 2024. 11

  23. [31]

    Physics of language models: Part 4.1, architecture design and the magic of canon layers

    Zeyuan Allen-Zhu. Physics of language models: Part 4.1, architecture design and the magic of canon layers. InProceedings of the 39th Conference on Neural Information Processing Systems, NeurIPS ’25, 2025. Full version available athttps://ssrn.com/abstract=5240330

  24. [33]

    David G. T. Barrett and Benoit Dherin. Implicit gradient regularization. InInternational Conference on Learning Representations, 2021

  25. [34]

    Smith, Benoit Dherin, David G

    Samuel L. Smith, Benoit Dherin, David G. T. Barrett, and Soham De. On the origin of implicit regularization in stochastic gradient descent. InInternational Conference on Learning Representations, 2021

  26. [35]

    On the trajectories of sgd without replacement.arXiv preprint arXiv:2312.16143, 2023

    Pierfrancesco Beneventano. On the trajectories of sgd without replacement.arXiv preprint arXiv:2312.16143, 2023

  27. [36]

    How neural networks learn the support is an implicit regularization effect of SGD

    Pierfrancesco Beneventano, Andrea Pinto, and Tomaso Poggio. How neural networks learn the support is an implicit regularization effect of SGD. 2024

  28. [37]

    Griffiths and J

    David F. Griffiths and J. M. Sanz-Serna. On the scope of the method of modified equations. SIAM Journal on Scientific and Statistical Computing, 7(3):994–1008, 1986

  29. [38]

    Springer, Berlin, Heidelberg, 2 edition, 2006

    Ernst Hairer, Christian Lubich, and Gerhard Wanner.Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations, volume 31 ofSpringer Series in Computational Mathematics. Springer, Berlin, Heidelberg, 2 edition, 2006

  30. [39]

    Implicit regularization in heavy- ball momentum accelerated stochastic gradient descent.arXiv preprint arXiv:2302.00849, 2023

    Avrajit Ghosh, He Lyu, Xitong Zhang, and Rongrong Wang. Implicit regularization in heavy- ball momentum accelerated stochastic gradient descent.arXiv preprint arXiv:2302.00849, 2023

  31. [40]

    Cattaneo, Jason Matthew Klusowski, and Boris Shigida

    Matias D. Cattaneo, Jason Matthew Klusowski, and Boris Shigida. On the implicit bias of Adam. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 5862–5906. PMLR, 2024

  32. [41]

    Implicit regularization in deep matrix factorization

    Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. Implicit regularization in deep matrix factorization. InAdvances in Neural Information Processing Systems, volume 32, 2019

  33. [42]

    Gradient descent converges linearly to flatter minima than gradient flow in shallow linear networks.arXiv preprint arXiv:2501.09137, 2025

    Pierfrancesco Beneventano and Blake Woodworth. Gradient descent converges linearly to flatter minima than gradient flow in shallow linear networks.arXiv preprint arXiv:2501.09137, 2025

  34. [43]

    Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A. Louis. Neural networks are a priori biased towards boolean functions with low entropy, 2020

  35. [44]

    Vapnik and Alexey Ya

    Vladimir N. Vapnik and Alexey Ya. Chervonenkis. On the uniform convergence of relative frequencies of events to their probabilities.Theory of Probability and Its Applications, 16(2):264–280, 1971

  36. [45]

    Vapnik.Statistical Learning Theory

    Vladimir N. Vapnik.Statistical Learning Theory. Wiley, 1998

  37. [46]

    Bartlett

    Peter L. Bartlett. The sample complexity of pattern classification with neural networks: The size of the weights is more important than the size of the network.IEEE Transactions on Information Theory, 44(2):525–536, 1998

  38. [47]

    Bartlett and Shahar Mendelson

    Peter L. Bartlett and Shahar Mendelson. Rademacher and gaussian complexities: Risk bounds and structural results.Journal of Machine Learning Research, 3:463–482, 2002

  39. [48]

    Norm-based capacity control in neural networks

    Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro. Norm-based capacity control in neural networks. InProceedings of The 28th Conference on Learning Theory, volume 40 of Proceedings of Machine Learning Research, pages 1376–1401. PMLR, 2015. 12

  40. [49]

    Bartlett, Dylan J

    Peter L. Bartlett, Dylan J. Foster, and Matus J. Telgarsky. Spectrally-normalized margin bounds for neural networks. InAdvances in Neural Information Processing Systems, volume 30, 2017

  41. [50]

    Understand- ing deep learning requires rethinking generalization

    Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understand- ing deep learning requires rethinking generalization. InInternational Conference on Learning Representations, 2017

  42. [51]

    Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien

    Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien. A closer look at memorization in deep networks. InProceedings of the 34th Internatio...

  43. [52]

    Edelman, Fred Zhang, and Boaz Barak

    Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, and Boaz Barak. SGD on neural networks learns functions of increasing complexity. InAdvances in Neural Information Processing Systems, volume 32, 2019

  44. [53]

    Train faster, generalize better: Stability of stochastic gradient descent

    Moritz Hardt, Benjamin Recht, and Yoram Singer. Train faster, generalize better: Stability of stochastic gradient descent. InProceedings of the 33rd International Conference on Machine Learning, volume 48 ofProceedings of Machine Learning Research, pages 1225–1234. PMLR, 2016

  45. [54]

    Exploring generalization in deep learning

    Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro. Exploring generalization in deep learning. InAdvances in Neural Information Processing Systems, volume 30, 2017

  46. [55]

    A PAC- bayesian approach to spectrally-normalized margin bounds for neural networks

    Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro. A PAC- bayesian approach to spectrally-normalized margin bounds for neural networks. InInterna- tional Conference on Learning Representations, 2018

  47. [56]

    Predicting the generalization gap in deep networks with margin distributions

    Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio. Predicting the generalization gap in deep networks with margin distributions. InInternational Conference on Learning Representations, 2019

  48. [57]

    Fantas- tic generalization measures and where to find them

    Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio. Fantas- tic generalization measures and where to find them. InInternational Conference on Learning Representations, 2020

  49. [58]

    Flat minima.Neural Computation, 9(1):1–42, 1997

    Sepp Hochreiter and Jürgen Schmidhuber. Flat minima.Neural Computation, 9(1):1–42, 1997

  50. [59]

    On large-batch training for deep learning: Generalization gap and sharp minima

    Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations, 2017

  51. [60]

    Sharp minima can generalize for deep nets

    Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets. InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 1019–1028. PMLR, 2017

  52. [61]

    Sharpness-aware minimization for efficiently improving generalization

    Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. InInternational Conference on Learning Representations, 2021

  53. [62]

    McAllester

    David A. McAllester. PAC-bayesian model averaging. InProceedings of the Twelfth Annual Conference on Computational Learning Theory, pages 164–170, 1999

  54. [63]

    Gintare Karolina Dziugaite and Daniel M. Roy. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. In Proceedings of the Thirty-Third Conference on Uncertainty in Artificial Intelligence, 2017

  55. [64]

    Adams, and Peter Orbanz

    Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz. Non- vacuous generalization bounds at the ImageNet scale: A PAC-bayesian compression approach. InInternational Conference on Learning Representations, 2019. 13

  56. [65]

    Stronger generalization bounds for deep nets via a compression approach

    Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. Stronger generalization bounds for deep nets via a compression approach. InInternational Conference on Learning Represen- tations, 2018

  57. [66]

    Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019

    Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019

  58. [67]

    Deep double descent: Where bigger models and more data hurt

    Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt. InInternational Conference on Learning Representations, 2020

  59. [68]

    Bartlett, Philip M

    Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler. Benign overfitting in linear regression.Proceedings of the National Academy of Sciences, 117(48):30063–30070, 2020

  60. [69]

    Neal.Bayesian Learning for Neural Networks, volume 118 ofLecture Notes in Statistics

    Radford M. Neal.Bayesian Learning for Neural Networks, volume 118 ofLecture Notes in Statistics. Springer, 1996

  61. [70]

    Christopher K. I. Williams. Computing with infinite networks. InAdvances in Neural Information Processing Systems, volume 9, 1996

  62. [71]

    Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein

    Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. Deep neural networks as gaussian processes. InInternational Conference on Learning Representations, 2018

  63. [72]

    Neural tangent kernel: Convergence and generalization in neural networks

    Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. InAdvances in Neural Information Processing Systems, volume 31, 2018

  64. [73]

    Camargo, and Ard A

    Kamaludin Dingle, Chico Q. Camargo, and Ard A. Louis. Input–output maps are strongly biased towards simple outputs.Nature Communications, 9:761, 2018

  65. [74]

    Random deep neural networks are biased towards simple functions

    Giacomo De Palma, Bobak Kiani, and Seth Lloyd. Random deep neural networks are biased towards simple functions. InAdvances in Neural Information Processing Systems, volume 32, 2019

  66. [75]

    On the complexity of finite sequences.IEEE Transactions on Information Theory, 22(1):75–81, 1976

    Abraham Lempel and Jacob Ziv. On the complexity of finite sequences.IEEE Transactions on Information Theory, 22(1):75–81, 1976

  67. [76]

    A universal algorithm for sequential data compression.IEEE Transactions on Information Theory, 23(3):337–343, 1977

    Jacob Ziv and Abraham Lempel. A universal algorithm for sequential data compression.IEEE Transactions on Information Theory, 23(3):337–343, 1977

  68. [77]

    Ming Li and Paul M. B. Vitányi.An Introduction to Kolmogorov Complexity and Its Applica- tions. Springer, 3 edition, 2008

  69. [78]

    Simplicity bias in transformers and their ability to learn sparse Boolean functions

    Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom. Simplicity bias in transformers and their ability to learn sparse Boolean functions. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5767–...

  70. [79]

    Michael Hahn and Mark Rofin. Why are sensitive functions hard for transformers? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14973–15008, Bangkok, Thailand, 2024. Association for Computational Linguistics

  71. [80]

    Transformers learn low sensitivity functions: Investigations and implications

    Bhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau, Youqi Huang, and Vatsal Sharan. Transformers learn low sensitivity functions: Investigations and implications. InInternational Conference on Learning Representations, 2025

  72. [81]

    Hamprecht, Yoshua Bengio, and Aaron Courville

    Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Le...

  73. [82]

    Abolafia, Jeffrey Pennington, and Jascha Sohl- Dickstein

    Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl- Dickstein. Sensitivity and generalization in neural networks: An empirical study. InInterna- tional Conference on Learning Representations, 2018

  74. [83]

    Complexity of linear regions in deep networks

    Boris Hanin and David Rolnick. Complexity of linear regions in deep networks. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 2596–2604. PMLR, 2019

  75. [84]

    Deep ReLU networks have surprisingly few activation patterns

    Boris Hanin and David Rolnick. Deep ReLU networks have surprisingly few activation patterns. InAdvances in Neural Information Processing Systems, volume 32, pages 359–368, 2019

  76. [85]

    Benoit Dherin, Michael Munn, Mihaela Rosca, and David G. T. Barrett. Why neural networks find simple solutions: The many regularizers of geometric complexity. InAdvances in Neural Information Processing Systems, volume 35, 2022

  77. [86]

    Neural networks trained with SGD learn distributions of increasing complexity

    Maria Refinetti, Alessandro Ingrosso, and Sebastian Goldt. Neural networks trained with SGD learn distributions of increasing complexity. InProceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 28843–...

  78. [87]

    Simplicity bias and optimization threshold in two-layer ReLU networks, 2024

    Etienne Boursier and Nicolas Flammarion. Simplicity bias and optimization threshold in two-layer ReLU networks, 2024

  79. [88]

    Saxe, and Peter E

    Yedi Zhang, Andrew M. Saxe, and Peter E. Latham. Saddle-to-saddle dynamics explains a simplicity bias across neural network architectures. InInternational Conference on Learning Representations, 2026. Poster

  80. [89]

    Saxe, James L

    Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. InInternational Conference on Learning Representations, 2014

  81. [90]

    Yaniv Blumenfeld, Dar Gilboa, and Daniel Soudry. Beyond signal propagation: Is feature diversity necessary in deep neural network initialization? InProceedings of the 37th Interna- tional Conference on Machine Learning, volume 119 ofProceedings of Machine Learning Research, pa...

  82. [91]

    Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2020

    Greg Yang. Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2020

  83. [92]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InProceedings of the 32nd International Conference on Machine Learning, volume 37 ofProceedings of Machine Learning Research, pages 448–456. PMLR, 2015

  84. [93]

    Theoretical analysis of auto rate-tuning by batch normalization

    Sanjeev Arora, Kaifeng Lyu, and Zhiyuan Li. Theoretical analysis of auto rate-tuning by batch normalization. InInternational Conference on Learning Representations, 2019

  85. [94]

    Reconciling modern deep learning with tradi- tional optimization analyses: The intrinsic learning rate

    Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora. Reconciling modern deep learning with tradi- tional optimization analyses: The intrinsic learning rate. InAdvances in Neural Information Processing Systems, volume 33, 2020

  86. [95]

    Number 32 in Notes on Applied Science

    James Hardy Wilkinson.Rounding Errors in Algebraic Processes. Number 32 in Notes on Applied Science. Her Majesty’s Stationery Office, London, 1963

  87. [96]

    Monographs on Numerical Analysis

    James Hardy Wilkinson.The Algebraic Eigenvalue Problem. Monographs on Numerical Analysis. Clarendon Press, Oxford, 1965

  88. [97]

    Higham.Accuracy and Stability of Numerical Algorithms

    Nicholas J. Higham.Accuracy and Stability of Numerical Algorithms. Society for Industrial and Applied Mathematics, Philadelphia, PA, 2 edition, 2002

  89. [98]

    M. P. Calvo, Ander Murua, and J. M. Sanz-Serna. Modified equations for ODEs. In Peter E. Kloeden and Kenneth J. Palmer, editors,Chaotic Numerics, volume 172 ofContemporary Mathematics, pages 63–74. American Mathematical Society, Providence, RI, 1994. 15

  90. [99]

    Modified equations for stochastic differential equations.BIT Numerical Mathematics, 46(1):111–125, 2006

    Tony Shardlow. Modified equations for stochastic differential equations.BIT Numerical Mathematics, 46(1):111–125, 2006

  91. [100]

    Zygalakis

    Konstantinos C. Zygalakis. On the existence and the applications of modified equations for stochastic differential equations.SIAM Journal on Scientific Computing, 33(1):102–130, 2011

  92. [101]

    Weak backward error analysis for SDEs.SIAM Journal on Numerical Analysis, 50(3):1735–1752, 2012

    Arnaud Debussche and Erwan Faou. Weak backward error analysis for SDEs.SIAM Journal on Numerical Analysis, 50(3):1735–1752, 2012

  93. [102]

    Uniform-in-time weak error analysis for stochastic gradient descent algorithms via diffusion approximation.Commu- nications in Mathematical Sciences, 18(1):163–188, 2020

    Yuanyuan Feng, Tingran Gao, Lei Li, Jian-Guo Liu, and Yulong Lu. Uniform-in-time weak error analysis for stochastic gradient descent algorithms via diffusion approximation.Commu- nications in Mathematical Sciences, 18(1):163–188, 2020

  94. [103]

    Stochastic modified equations and adaptive stochastic gradient algorithms

    Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and adaptive stochastic gradient algorithms. InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 2101–2110. PMLR, 2017

  95. [104]

    Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations.Journal of Machine Learning Research, 20(40):1–47, 2019

    Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations.Journal of Machine Learning Research, 20(40):1–47, 2019

  96. [105]

    Toward equation of motion for deep neural networks: Continuous-time gra- dient descent and discretization error analysis

    Taiki Miyagawa. Toward equation of motion for deep neural networks: Continuous-time gra- dient descent and discretization error analysis. InAdvances in Neural Information Processing Systems, volume 35, pages 37778–37791, 2022

  97. [106]

    On a continuous time model of gradient descent dynamics and instability in deep learning.arXiv preprint arXiv:2302.01952, 2023

    Mihaela Rosca, Yan Wu, Chongli Qin, and Benoit Dherin. On a continuous time model of gradient descent dynamics and instability in deep learning.arXiv preprint arXiv:2302.01952, 2023

  98. [107]

    Modified loss of momentum gradient descent: Fine- grained analysis.arXiv preprint arXiv:2509.08483, 2025

    Matias D Cattaneo and Boris Shigida. Modified loss of momentum gradient descent: Fine- grained analysis.arXiv preprint arXiv:2509.08483, 2025

  99. [108]

    Higham, and Konstantinos C

    Stefano Di Giovacchino, Desmond J. Higham, and Konstantinos C. Zygalakis. Backward error analysis and the qualitative behaviour of stochastic optimization algorithms: Application to stochastic coordinate descent.Journal of Computational Dynamics, 11(4):453–467, 2024

  100. [109]

    How memory in optimization algorithms implicitly modifies the loss.Advances in Neural Information Processing Systems, 38:156059–156096, 2026

    Matias Cattaneo and Boris Shigida. How memory in optimization algorithms implicitly modifies the loss.Advances in Neural Information Processing Systems, 38:156059–156096, 2026

  101. [110]

    The effect of mini-batch noise on the implicit bias of adam.arXiv preprint arXiv:2602.01642, 2026

    Matias D Cattaneo and Boris Shigida. The effect of mini-batch noise on the implicit bias of adam.arXiv preprint arXiv:2602.01642, 2026

  102. [111]

    The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19(70):1–57, 2018

    Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19(70):1–57, 2018

  103. [112]

    Lee, Daniel Soudry, and Nathan Srebro

    Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nathan Srebro. Implicit bias of gradient descent on linear convolutional networks. InAdvances in Neural Information Processing Systems, volume 31, 2018

  104. [113]

    Gradient descent maximizes the margin of homogeneous neural networks

    Kaifeng Lyu and Jian Li. Gradient descent maximizes the margin of homogeneous neural networks. InInternational Conference on Learning Representations, 2020

  105. [114]

    Train longer, generalize better: Closing the gen- eralization gap in large batch training of neural networks

    Elad Hoffer, Itay Hubara, and Daniel Soudry. Train longer, generalize better: Closing the gen- eralization gap in large batch training of neural networks. InAdvances in Neural Information Processing Systems, volume 30, 2017

  106. [115]

    Hoffman, and David M

    Stephan Mandt, Matthew D. Hoffman, and David M. Blei. Stochastic gradient descent as approximate bayesian inference.Journal of Machine Learning Research, 18(134):1–35, 2017

  107. [116]

    Spread” column reports maxσw acc−min σw acc (in percentage points); “Train

    Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht. The marginal value of adaptive gradient methods in machine learning. InAdvances in Neural Information Processing Systems, volume 30, 2017. 16 A Experimental Details This appendix records the s...

  108. [117]

    Our results are consistent with this broader message: Adam-family methods do not merely train faster in our grid; they erase initialization-scale dependence more readily

    showed that adaptive methods can select different solutions from SGD and can generalize differently even when they optimize training loss well. Our results are consistent with this broader message: Adam-family methods do not merely train faster in our grid; they erase initiali...

  109. [118]

    go further, showing that models can retain fingerprints of their training seed. These papers motivate the same broad question as ours in a different setting: which parts of final performance are due to the initial condition, and which are erased by training? Controlled languag...

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.