REVIEW 2 major objections 1 minor 117 references
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read The inductive bias of a trained neural network is its architectural prior filtered by the forgetting dynamics of the training pipeline.
desk verdict Low-LR SGD keeps a 26-point test-acc spread across init scales on CIFAR ResNets even at 99.5% train acc, while Adam and L2 erase it, but the spread may trace to other trajectory differences rather than retained initialization memory. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Initialization memory, the dependence of the validation-selected predictor on the scale of the random initialization; it quantifies how much initial bias survives training.
What would settle it
An experiment in which the accuracy spread across init scales disappears when other optimization factors are controlled while maintaining the same training accuracy.
Extended reading notes
Core claim
In controlled experiments, initialization memory—the dependence of the final predictor on random initialization scale—persists under gradient-flow-like dynamics such as low-LR SGD but is erased on timescales set by stochastic effects, norm decay, or adaptive preconditioning; therefore the practical bias equals the architectural prior after filtering by forgetting dynamics, and regularizers improve generalization precisely by erasing initialization memory.
Load-bearing premise
Variation in test accuracy across initialization scales after high training accuracy isolates retained initialization memory rather than other differences in optimization.
Editorial extensions
If this is right
- Low-learning-rate SGD interpolates yet retains initialization memory, leading to large test accuracy variation.
- Adam-family methods erase the dependence on initialization scale.
- Pairing larger learning rates with L2 norm control causes SGD to forget initialization.
- The time scale of forgetting is governed by the size of explicit or implicit regularization.
Reading between the lines
- Initialization scale may need to be tuned differently depending on the optimizer used.
- This forgetting view could explain why certain training choices improve generalization beyond what architecture alone predicts.
- Extending training time does not necessarily increase forgetting if the regime preserves memory.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that trained deep networks retain a measurable dependence on the scale of their random initialization ('initialization memory'), which survives low-learning-rate SGD but is erased by Adam-family optimizers or explicit L2 regularization. This is demonstrated via controlled CIFAR-10 ResNet experiments showing a 26.5 percentage point spread in test accuracy across initialization scales at >=99.5% training accuracy; the spread persists even after extending training to 5000 epochs. The authors interpret the results as evidence that practical inductive bias is the architectural prior filtered by the forgetting dynamics of the training pipeline, with the same regularizers that aid generalization also erasing initialization memory.
Significance. If the reported accuracy spreads can be isolated to retained initialization memory rather than correlated differences in optimization trajectories, the work offers a useful empirical lens on how training dynamics shape effective priors. The controlled regime comparisons (low-LR SGD vs. Adam vs. L2) and the observation that extended epochs do not close the gap provide concrete, falsifiable distinctions between regimes. The absence of free parameters or fitted models in the core measurements is a strength of the empirical design.
major comments (2)
- [Abstract] Abstract and the CIFAR-10 ResNet-9 experiments: the central attribution of the 26.5 pp test-accuracy spread to retained 'initialization memory' (rather than other init-scale-dependent trajectory factors such as effective step-size distribution, basin geometry, or final weight norms) is load-bearing for the forgetting-time interpretation, yet the reported controls (high training accuracy plus extension to 5000 epochs) do not constrain these alternatives. High train accuracy plus fixed epoch count is insufficient to isolate the claimed mechanism.
- [Abstract] The manuscript reports quantitative separation under controlled setups but provides no details on the number of independent runs, error bars, or exact hyperparameter sweep ranges for the 26.5 pp spread. Without these, it is impossible to assess whether the observed variation is statistically robust or sensitive to minor implementation choices.
minor comments (1)
- Notation for 'initialization memory' is introduced without a formal definition or equation; a precise mathematical statement of the dependence being measured would improve clarity.
Simulated Author's Rebuttal
We thank the referee for the careful reading and constructive feedback. We respond to each major comment below.
read point-by-point responses
-
Referee: [Abstract] Abstract and the CIFAR-10 ResNet-9 experiments: the central attribution of the 26.5 pp test-accuracy spread to retained 'initialization memory' (rather than other init-scale-dependent trajectory factors such as effective step-size distribution, basin geometry, or final weight norms) is load-bearing for the forgetting-time interpretation, yet the reported controls (high training accuracy plus extension to 5000 epochs) do not constrain these alternatives. High train accuracy plus fixed epoch count is insufficient to isolate the claimed mechanism.
Authors: We define initialization memory operationally as the dependence of the validation-selected predictor on initialization scale under a fixed training procedure. The reported experiments show that this dependence survives low-LR SGD even after ≥99.5% training accuracy is reached and after training is extended to 5000 epochs, while the same dependence is erased under Adam or explicit L2 regularization. These regime contrasts are the primary evidence for the forgetting-time view. We agree that additional measurements (e.g., final weight norms across scales or basin geometry) would further constrain alternative explanations and will add a limitations paragraph discussing this point in the revision. revision: partial
-
Referee: [Abstract] The manuscript reports quantitative separation under controlled setups but provides no details on the number of independent runs, error bars, or exact hyperparameter sweep ranges for the 26.5 pp spread. Without these, it is impossible to assess whether the observed variation is statistically robust or sensitive to minor implementation choices.
Authors: We agree that these experimental details are necessary. The 26.5 pp figure is computed from multiple independent runs that vary only the initialization scale (different random seeds), and we will include the exact number of runs, standard-error bars, and the precise hyperparameter ranges used for all reported regimes in the revised manuscript and supplementary material. revision: yes
Circularity Check
No circularity; purely empirical measurements with no self-referential derivations
full rationale
The paper reports controlled CIFAR-10 experiments on ResNets that measure test-accuracy spread across initialization scales under fixed training-accuracy thresholds. No equations, fitted parameters, or predictions are defined in terms of the target quantities. No self-citations are invoked as load-bearing uniqueness theorems or ansatzes. The central observations (e.g., 26.5 pp spread under low-LR SGD) are direct empirical outputs, not reductions of any input by construction. The work is therefore self-contained against external benchmarks.
Assumptions & free parameters
assumptions (1)
- domain assumption Standard assumptions underlying SGD and Adam dynamics on non-convex loss surfaces
invented entities (1)
-
initialization memory
Cite this review
Pith. "Pith review of Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias." pith.science (2026). https://pith.science/paper/MLZBYQUY
@misc{pith2026260529152,
author = {Pith},
title = {Pith review of: Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias},
year = {2026},
howpublished = {\url{https://pith.science/paper/MLZBYQUY}},
note = {Machine review of arXiv:2605.29152}
}
abstract
Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survives the training pipeline. To make the question measurable, we introduce initialization memory: the dependence of the validation-selected predictor on the scale of the random initialization. We perform controlled CIFAR-10 experiments on ResNets where initialization memory already sharply separates training regimes. Low-learning-rate SGD can interpolate while still remembering its initialization: on ResNet-9 with batch size $b=128$, test accuracy varies by $26.5$ percentage points across initialization scales despite $\ge99.5\%$ training accuracy. This is not undertraining: extending the same low-learning-rate regime to $5{,}000$ epochs leaves the spread essentially unchanged. In contrast, Adam-family methods largely erase the dependence. SGD can also be made to forget when larger learning rates are paired with explicit $L_2$ norm control. We interpret these findings in terms of the time scale of forgetting: gradient-flow-like dynamics can preserve initialization memory, whereas stochastic finite-step effects, explicit norm decay, and adaptive preconditioning erase it on scales governed by the size of explicit or implicit regularization. The practical inductive bias of a trained network is therefore not the architectural prior alone, but the architectural prior after being filtered by the forgetting dynamics of the training pipeline; and the same regularizers that improve generalization are precisely those that erase memory of initialization.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Under- standing deep learning (still) requires rethinking generalization.Communications of the ACM, 64(3):107–115, 2021
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Under- standing deep learning (still) requires rethinking generalization.Communications of the ACM, 64(3):107–115, 2021
2021
-
[2]
Strogatz.Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering
Steven H. Strogatz.Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering. Westview Press, 2 edition, 2015
2015
-
[3]
Hirsch, Stephen Smale, and Robert L
Morris W. Hirsch, Stephen Smale, and Robert L. Devaney.Differential Equations, Dynamical Systems, and an Introduction to Chaos. Academic Press, 3 edition, 2013
2013
-
[4]
On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks
Hancheng Min, Salma Tarmoun, René Vidal, and Enrique Mallada. On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks. InProceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 7760–7768. PMLR, 2021
2021
-
[5]
On the role of initialization on the implicit bias in deep linear networks, 2024
Oria Gruber and Haim Avron. On the role of initialization on the implicit bias in deep linear networks, 2024
2024
-
[6]
Camargo, and Ard A
Guillermo Valle-Pérez, Chico Q. Camargo, and Ard A. Louis. Deep learning generalizes because the parameter–function map is biased towards simple functions. InInternational Conference on Learning Representations, 2019
2019
-
[7]
Chris Mingard, Henry Rees, Guillermo Valle-Pérez, and Ard A. Louis. Deep neural networks have an inbuilt occam’s razor.Nature Communications, 16:220, 2025
2025
-
[8]
Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026
Thomas Fink. Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026
Show all 117 references
-
[9]
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. InProceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, volume 9 ofProceedings of Machine Learning Research, pages 249–256. P...
2010
-
[10]
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. InProceedings of the IEEE International Conference on Computer Vision, pages 1026–1034, 2015
2015
-
[11]
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli. Exponential expressivity in deep neural networks through transient chaos. InAdvances in Neural Information Processing Systems, volume 29, 2016
2016
-
[12]
Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. Deep information propagation. InInternational Conference on Learning Representations, 2017
2017
-
[13]
Schoenholz, and Surya Ganguli
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli. Resurrecting the sigmoid in deep learning through dynamical isometry: Theory and practice. InAdvances in Neural Information Processing Systems, volume 30, 2017
2017
-
[14]
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick. How to start training: The effect of initialization and architecture. InAdvances in Neural Information Processing Systems, volume 31, 2018
2018
-
[15]
Which neural net architectures give rise to exploding and vanishing gradients? InAdvances in Neural Information Processing Systems, volume 31, 2018
Boris Hanin. Which neural net architectures give rise to exploding and vanishing gradients? InAdvances in Neural Information Processing Systems, volume 31, 2018. 10
2018
-
[16]
Schoenholz, and Jeffrey Pennington
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S. Schoenholz, and Jeffrey Pennington. Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks. InProceedings of the 35th International Conference on Machine L...
2018
-
[17]
Schoenholz
Minmin Chen, Jeffrey Pennington, and Samuel S. Schoenholz. Dynamical isometry and a mean field theory of RNNs: Gating enables signal propagation in recurrent neural networks. InProceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Ma...
2018
-
[18]
Greg Yang and Edward J. Hu. Tensor programs IV: Feature learning in infinite-width neural networks. InProceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 11727–11737. PMLR, 2021
2021
-
[19]
Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor programs V: Tuning large neural networks via zero-shot hyperparameter transfer. InAdvances in Neural Information Processing Syst...
2021
-
[20]
Self-consistent dynamical field theory of kernel evo- lution in wide neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2023(11):114009, 2023
Blake Bordelon and Cengiz Pehlevan. Self-consistent dynamical field theory of kernel evo- lution in wide neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2023(11):114009, 2023
2023
-
[21]
Dynamics of finite width kernel and prediction fluctua- tions in mean field neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2024(10):104021, 2024
Blake Bordelon and Cengiz Pehlevan. Dynamics of finite width kernel and prediction fluctua- tions in mean field neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2024(10):104021, 2024
2024
-
[22]
Deep linear network training dynamics from random initialization: Data, width, depth, and hyperparameter transfer
Blake Bordelon and Cengiz Pehlevan. Deep linear network training dynamics from random initialization: Data, width, depth, and hyperparameter transfer. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceedings of Machine Learning Research,...
2025
-
[23]
Adaptive kernel predictors from feature-learning infinite limits of neural networks
Clarissa Lauditi, Blake Bordelon, and Cengiz Pehlevan. Adaptive kernel predictors from feature-learning infinite limits of neural networks. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceedings of Machine Learning Research, pages 3261...
2025
-
[24]
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping.CoRR, abs/2002.06305, 2020
2002
-
[25]
Zuidema, and Stella R
Oskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra, Hailey Schoelkopf, Willem H. Zuidema, and Stella R. Biderman. PolyPythias: Stability and outliers across fifty language model pre-training runs. InInternational Conference on Learning Representations, 2025
2025
-
[26]
Convergence and divergence of language models under different random seeds
Finlay Fehlauer, Kyle Mahowald, and Tiago Pimentel. Convergence and divergence of language models under different random seeds. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32982–32991, Suzhou, China,
2025
-
[27]
Association for Computational Linguistics
-
[28]
SeedPrints: Fingerprints can even tell which seed your large language model was trained from
Yao Tong, Haonan Wang, Siquan Li, Kenji Kawaguchi, and Tianyang Hu. SeedPrints: Fingerprints can even tell which seed your large language model was trained from. In International Conference on Learning Representations, 2026. Poster
2026
-
[29]
Transformers are born biased: Structural inductive biases at random initialization and their practical consequences, 2026
Siquan Li, Yao Tong, Haonan Wang, and Tianyang Hu. Transformers are born biased: Structural inductive biases at random initialization and their practical consequences, 2026
2026
-
[30]
Physics of language models: Part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 1067–1077. PMLR, 2024. 11
2024
-
[31]
Physics of language models: Part 4.1, architecture design and the magic of canon layers
Zeyuan Allen-Zhu. Physics of language models: Part 4.1, architecture design and the magic of canon layers. InProceedings of the 39th Conference on Neural Information Processing Systems, NeurIPS ’25, 2025. Full version available athttps://ssrn.com/abstract=5240330
2025
-
[33]
David G. T. Barrett and Benoit Dherin. Implicit gradient regularization. InInternational Conference on Learning Representations, 2021
2021
-
[34]
Smith, Benoit Dherin, David G
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, and Soham De. On the origin of implicit regularization in stochastic gradient descent. InInternational Conference on Learning Representations, 2021
2021
-
[35]
On the trajectories of sgd without replacement.arXiv preprint arXiv:2312.16143, 2023
Pierfrancesco Beneventano. On the trajectories of sgd without replacement.arXiv preprint arXiv:2312.16143, 2023
2023
-
[36]
How neural networks learn the support is an implicit regularization effect of SGD
Pierfrancesco Beneventano, Andrea Pinto, and Tomaso Poggio. How neural networks learn the support is an implicit regularization effect of SGD. 2024
2024
-
[37]
Griffiths and J
David F. Griffiths and J. M. Sanz-Serna. On the scope of the method of modified equations. SIAM Journal on Scientific and Statistical Computing, 7(3):994–1008, 1986
1986
-
[38]
Springer, Berlin, Heidelberg, 2 edition, 2006
Ernst Hairer, Christian Lubich, and Gerhard Wanner.Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations, volume 31 ofSpringer Series in Computational Mathematics. Springer, Berlin, Heidelberg, 2 edition, 2006
2006
-
[39]
Implicit regularization in heavy- ball momentum accelerated stochastic gradient descent.arXiv preprint arXiv:2302.00849, 2023
Avrajit Ghosh, He Lyu, Xitong Zhang, and Rongrong Wang. Implicit regularization in heavy- ball momentum accelerated stochastic gradient descent.arXiv preprint arXiv:2302.00849, 2023
2023
-
[40]
Cattaneo, Jason Matthew Klusowski, and Boris Shigida
Matias D. Cattaneo, Jason Matthew Klusowski, and Boris Shigida. On the implicit bias of Adam. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 5862–5906. PMLR, 2024
2024
-
[41]
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. Implicit regularization in deep matrix factorization. InAdvances in Neural Information Processing Systems, volume 32, 2019
2019
-
[42]
Gradient descent converges linearly to flatter minima than gradient flow in shallow linear networks.arXiv preprint arXiv:2501.09137, 2025
Pierfrancesco Beneventano and Blake Woodworth. Gradient descent converges linearly to flatter minima than gradient flow in shallow linear networks.arXiv preprint arXiv:2501.09137, 2025
2025
-
[43]
Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A. Louis. Neural networks are a priori biased towards boolean functions with low entropy, 2020
2020
-
[44]
Vapnik and Alexey Ya
Vladimir N. Vapnik and Alexey Ya. Chervonenkis. On the uniform convergence of relative frequencies of events to their probabilities.Theory of Probability and Its Applications, 16(2):264–280, 1971
1971
-
[45]
Vapnik.Statistical Learning Theory
Vladimir N. Vapnik.Statistical Learning Theory. Wiley, 1998
1998
-
[46]
Bartlett
Peter L. Bartlett. The sample complexity of pattern classification with neural networks: The size of the weights is more important than the size of the network.IEEE Transactions on Information Theory, 44(2):525–536, 1998
1998
-
[47]
Bartlett and Shahar Mendelson
Peter L. Bartlett and Shahar Mendelson. Rademacher and gaussian complexities: Risk bounds and structural results.Journal of Machine Learning Research, 3:463–482, 2002
2002
-
[48]
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro. Norm-based capacity control in neural networks. InProceedings of The 28th Conference on Learning Theory, volume 40 of Proceedings of Machine Learning Research, pages 1376–1401. PMLR, 2015. 12
2015
-
[49]
Bartlett, Dylan J
Peter L. Bartlett, Dylan J. Foster, and Matus J. Telgarsky. Spectrally-normalized margin bounds for neural networks. InAdvances in Neural Information Processing Systems, volume 30, 2017
2017
-
[50]
Understand- ing deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understand- ing deep learning requires rethinking generalization. InInternational Conference on Learning Representations, 2017
2017
-
[51]
Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien. A closer look at memorization in deep networks. InProceedings of the 34th Internatio...
2017
-
[52]
Edelman, Fred Zhang, and Boaz Barak
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, and Boaz Barak. SGD on neural networks learns functions of increasing complexity. InAdvances in Neural Information Processing Systems, volume 32, 2019
2019
-
[53]
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer. Train faster, generalize better: Stability of stochastic gradient descent. InProceedings of the 33rd International Conference on Machine Learning, volume 48 ofProceedings of Machine Learning Research, pages 1225–1234. PMLR, 2016
2016
-
[54]
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro. Exploring generalization in deep learning. InAdvances in Neural Information Processing Systems, volume 30, 2017
2017
-
[55]
A PAC- bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro. A PAC- bayesian approach to spectrally-normalized margin bounds for neural networks. InInterna- tional Conference on Learning Representations, 2018
2018
-
[56]
Predicting the generalization gap in deep networks with margin distributions
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio. Predicting the generalization gap in deep networks with margin distributions. InInternational Conference on Learning Representations, 2019
2019
-
[57]
Fantas- tic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio. Fantas- tic generalization measures and where to find them. InInternational Conference on Learning Representations, 2020
2020
-
[58]
Flat minima.Neural Computation, 9(1):1–42, 1997
Sepp Hochreiter and Jürgen Schmidhuber. Flat minima.Neural Computation, 9(1):1–42, 1997
1997
-
[59]
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations, 2017
2017
-
[60]
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets. InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 1019–1028. PMLR, 2017
2017
-
[61]
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. InInternational Conference on Learning Representations, 2021
2021
-
[62]
McAllester
David A. McAllester. PAC-bayesian model averaging. InProceedings of the Twelfth Annual Conference on Computational Learning Theory, pages 164–170, 1999
1999
-
[63]
Gintare Karolina Dziugaite and Daniel M. Roy. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. In Proceedings of the Thirty-Third Conference on Uncertainty in Artificial Intelligence, 2017
2017
-
[64]
Adams, and Peter Orbanz
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz. Non- vacuous generalization bounds at the ImageNet scale: A PAC-bayesian compression approach. InInternational Conference on Learning Representations, 2019. 13
2019
-
[65]
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. Stronger generalization bounds for deep nets via a compression approach. InInternational Conference on Learning Represen- tations, 2018
2018
-
[66]
Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019
2019
-
[67]
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt. InInternational Conference on Learning Representations, 2020
2020
-
[68]
Bartlett, Philip M
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler. Benign overfitting in linear regression.Proceedings of the National Academy of Sciences, 117(48):30063–30070, 2020
2020
-
[69]
Neal.Bayesian Learning for Neural Networks, volume 118 ofLecture Notes in Statistics
Radford M. Neal.Bayesian Learning for Neural Networks, volume 118 ofLecture Notes in Statistics. Springer, 1996
1996
-
[70]
Christopher K. I. Williams. Computing with infinite networks. InAdvances in Neural Information Processing Systems, volume 9, 1996
1996
-
[71]
Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. Deep neural networks as gaussian processes. InInternational Conference on Learning Representations, 2018
2018
-
[72]
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. InAdvances in Neural Information Processing Systems, volume 31, 2018
2018
-
[73]
Camargo, and Ard A
Kamaludin Dingle, Chico Q. Camargo, and Ard A. Louis. Input–output maps are strongly biased towards simple outputs.Nature Communications, 9:761, 2018
2018
-
[74]
Random deep neural networks are biased towards simple functions
Giacomo De Palma, Bobak Kiani, and Seth Lloyd. Random deep neural networks are biased towards simple functions. InAdvances in Neural Information Processing Systems, volume 32, 2019
2019
-
[75]
On the complexity of finite sequences.IEEE Transactions on Information Theory, 22(1):75–81, 1976
Abraham Lempel and Jacob Ziv. On the complexity of finite sequences.IEEE Transactions on Information Theory, 22(1):75–81, 1976
1976
-
[76]
A universal algorithm for sequential data compression.IEEE Transactions on Information Theory, 23(3):337–343, 1977
Jacob Ziv and Abraham Lempel. A universal algorithm for sequential data compression.IEEE Transactions on Information Theory, 23(3):337–343, 1977
1977
-
[77]
Ming Li and Paul M. B. Vitányi.An Introduction to Kolmogorov Complexity and Its Applica- tions. Springer, 3 edition, 2008
2008
-
[78]
Simplicity bias in transformers and their ability to learn sparse Boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom. Simplicity bias in transformers and their ability to learn sparse Boolean functions. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5767–...
2023
-
[79]
Michael Hahn and Mark Rofin. Why are sensitive functions hard for transformers? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14973–15008, Bangkok, Thailand, 2024. Association for Computational Linguistics
2024
-
[80]
Transformers learn low sensitivity functions: Investigations and implications
Bhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau, Youqi Huang, and Vatsal Sharan. Transformers learn low sensitivity functions: Investigations and implications. InInternational Conference on Learning Representations, 2025
2025
-
[81]
Hamprecht, Yoshua Bengio, and Aaron Courville
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Le...
2019
-
[82]
Abolafia, Jeffrey Pennington, and Jascha Sohl- Dickstein
Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl- Dickstein. Sensitivity and generalization in neural networks: An empirical study. InInterna- tional Conference on Learning Representations, 2018
2018
-
[83]
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick. Complexity of linear regions in deep networks. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 2596–2604. PMLR, 2019
2019
-
[84]
Deep ReLU networks have surprisingly few activation patterns
Boris Hanin and David Rolnick. Deep ReLU networks have surprisingly few activation patterns. InAdvances in Neural Information Processing Systems, volume 32, pages 359–368, 2019
2019
-
[85]
Benoit Dherin, Michael Munn, Mihaela Rosca, and David G. T. Barrett. Why neural networks find simple solutions: The many regularizers of geometric complexity. InAdvances in Neural Information Processing Systems, volume 35, 2022
2022
-
[86]
Neural networks trained with SGD learn distributions of increasing complexity
Maria Refinetti, Alessandro Ingrosso, and Sebastian Goldt. Neural networks trained with SGD learn distributions of increasing complexity. InProceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 28843–...
2023
-
[87]
Simplicity bias and optimization threshold in two-layer ReLU networks, 2024
Etienne Boursier and Nicolas Flammarion. Simplicity bias and optimization threshold in two-layer ReLU networks, 2024
2024
-
[88]
Saxe, and Peter E
Yedi Zhang, Andrew M. Saxe, and Peter E. Latham. Saddle-to-saddle dynamics explains a simplicity bias across neural network architectures. InInternational Conference on Learning Representations, 2026. Poster
2026
-
[89]
Saxe, James L
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. InInternational Conference on Learning Representations, 2014
2014
-
[90]
Yaniv Blumenfeld, Dar Gilboa, and Daniel Soudry. Beyond signal propagation: Is feature diversity necessary in deep neural network initialization? InProceedings of the 37th Interna- tional Conference on Machine Learning, volume 119 ofProceedings of Machine Learning Research, pa...
2020
-
[91]
Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2020
Greg Yang. Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2020
2020
-
[92]
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InProceedings of the 32nd International Conference on Machine Learning, volume 37 ofProceedings of Machine Learning Research, pages 448–456. PMLR, 2015
2015
-
[93]
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Kaifeng Lyu, and Zhiyuan Li. Theoretical analysis of auto rate-tuning by batch normalization. InInternational Conference on Learning Representations, 2019
2019
-
[94]
Reconciling modern deep learning with tradi- tional optimization analyses: The intrinsic learning rate
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora. Reconciling modern deep learning with tradi- tional optimization analyses: The intrinsic learning rate. InAdvances in Neural Information Processing Systems, volume 33, 2020
2020
-
[95]
Number 32 in Notes on Applied Science
James Hardy Wilkinson.Rounding Errors in Algebraic Processes. Number 32 in Notes on Applied Science. Her Majesty’s Stationery Office, London, 1963
1963
-
[96]
Monographs on Numerical Analysis
James Hardy Wilkinson.The Algebraic Eigenvalue Problem. Monographs on Numerical Analysis. Clarendon Press, Oxford, 1965
1965
-
[97]
Higham.Accuracy and Stability of Numerical Algorithms
Nicholas J. Higham.Accuracy and Stability of Numerical Algorithms. Society for Industrial and Applied Mathematics, Philadelphia, PA, 2 edition, 2002
2002
-
[98]
M. P. Calvo, Ander Murua, and J. M. Sanz-Serna. Modified equations for ODEs. In Peter E. Kloeden and Kenneth J. Palmer, editors,Chaotic Numerics, volume 172 ofContemporary Mathematics, pages 63–74. American Mathematical Society, Providence, RI, 1994. 15
1994
-
[99]
Modified equations for stochastic differential equations.BIT Numerical Mathematics, 46(1):111–125, 2006
Tony Shardlow. Modified equations for stochastic differential equations.BIT Numerical Mathematics, 46(1):111–125, 2006
2006
-
[100]
Zygalakis
Konstantinos C. Zygalakis. On the existence and the applications of modified equations for stochastic differential equations.SIAM Journal on Scientific Computing, 33(1):102–130, 2011
2011
-
[101]
Weak backward error analysis for SDEs.SIAM Journal on Numerical Analysis, 50(3):1735–1752, 2012
Arnaud Debussche and Erwan Faou. Weak backward error analysis for SDEs.SIAM Journal on Numerical Analysis, 50(3):1735–1752, 2012
2012
-
[102]
Uniform-in-time weak error analysis for stochastic gradient descent algorithms via diffusion approximation.Commu- nications in Mathematical Sciences, 18(1):163–188, 2020
Yuanyuan Feng, Tingran Gao, Lei Li, Jian-Guo Liu, and Yulong Lu. Uniform-in-time weak error analysis for stochastic gradient descent algorithms via diffusion approximation.Commu- nications in Mathematical Sciences, 18(1):163–188, 2020
2020
-
[103]
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and adaptive stochastic gradient algorithms. InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 2101–2110. PMLR, 2017
2017
-
[104]
Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations.Journal of Machine Learning Research, 20(40):1–47, 2019
Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations.Journal of Machine Learning Research, 20(40):1–47, 2019
2019
-
[105]
Toward equation of motion for deep neural networks: Continuous-time gra- dient descent and discretization error analysis
Taiki Miyagawa. Toward equation of motion for deep neural networks: Continuous-time gra- dient descent and discretization error analysis. InAdvances in Neural Information Processing Systems, volume 35, pages 37778–37791, 2022
2022
-
[106]
On a continuous time model of gradient descent dynamics and instability in deep learning.arXiv preprint arXiv:2302.01952, 2023
Mihaela Rosca, Yan Wu, Chongli Qin, and Benoit Dherin. On a continuous time model of gradient descent dynamics and instability in deep learning.arXiv preprint arXiv:2302.01952, 2023
2023
-
[107]
Modified loss of momentum gradient descent: Fine- grained analysis.arXiv preprint arXiv:2509.08483, 2025
Matias D Cattaneo and Boris Shigida. Modified loss of momentum gradient descent: Fine- grained analysis.arXiv preprint arXiv:2509.08483, 2025
2025
-
[108]
Higham, and Konstantinos C
Stefano Di Giovacchino, Desmond J. Higham, and Konstantinos C. Zygalakis. Backward error analysis and the qualitative behaviour of stochastic optimization algorithms: Application to stochastic coordinate descent.Journal of Computational Dynamics, 11(4):453–467, 2024
2024
-
[109]
How memory in optimization algorithms implicitly modifies the loss.Advances in Neural Information Processing Systems, 38:156059–156096, 2026
Matias Cattaneo and Boris Shigida. How memory in optimization algorithms implicitly modifies the loss.Advances in Neural Information Processing Systems, 38:156059–156096, 2026
2026
-
[110]
The effect of mini-batch noise on the implicit bias of adam.arXiv preprint arXiv:2602.01642, 2026
Matias D Cattaneo and Boris Shigida. The effect of mini-batch noise on the implicit bias of adam.arXiv preprint arXiv:2602.01642, 2026
2026 arXiv
-
[111]
The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19(70):1–57, 2018
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19(70):1–57, 2018
2018
-
[112]
Lee, Daniel Soudry, and Nathan Srebro
Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nathan Srebro. Implicit bias of gradient descent on linear convolutional networks. InAdvances in Neural Information Processing Systems, volume 31, 2018
2018
-
[113]
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li. Gradient descent maximizes the margin of homogeneous neural networks. InInternational Conference on Learning Representations, 2020
2020
-
[114]
Train longer, generalize better: Closing the gen- eralization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry. Train longer, generalize better: Closing the gen- eralization gap in large batch training of neural networks. InAdvances in Neural Information Processing Systems, volume 30, 2017
2017
-
[115]
Hoffman, and David M
Stephan Mandt, Matthew D. Hoffman, and David M. Blei. Stochastic gradient descent as approximate bayesian inference.Journal of Machine Learning Research, 18(134):1–35, 2017
2017
-
[116]
Spread” column reports maxσw acc−min σw acc (in percentage points); “Train
Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht. The marginal value of adaptive gradient methods in machine learning. InAdvances in Neural Information Processing Systems, volume 30, 2017. 16 A Experimental Details This appendix records the s...
2017
-
[117]
Our results are consistent with this broader message: Adam-family methods do not merely train faster in our grid; they erase initialization-scale dependence more readily
showed that adaptive methods can select different solutions from SGD and can generalize differently even when they optimize training loss well. Our results are consistent with this broader message: Adam-family methods do not merely train faster in our grid; they erase initiali...
-
[118]
go further, showing that models can retain fingerprints of their training seed. These papers motivate the same broad question as ours in a different setting: which parts of final performance are due to the initial condition, and which are erased by training? Controlled languag...
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.