Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
arXiv preprint arXiv:2101.12176 , year=
6 Pith papers cite this work. Polarity classification is still indexing.
abstract
For infinitesimal learning rates, stochastic gradient descent (SGD) follows the path of gradient flow on the full batch loss function. However moderately large learning rates can achieve higher test accuracies, and this generalization benefit is not explained by convergence bounds, since the learning rate which maximizes test accuracy is often larger than the learning rate which minimizes training loss. To interpret this phenomenon we prove that for SGD with random shuffling, the mean SGD iterate also stays close to the path of gradient flow if the learning rate is small and finite, but on a modified loss. This modified loss is composed of the original loss function and an implicit regularizer, which penalizes the norms of the minibatch gradients. Under mild assumptions, when the batch size is small the scale of the implicit regularization term is proportional to the ratio of the learning rate to the batch size. We verify empirically that explicitly including the implicit regularizer in the loss can enhance the test accuracy when the learning rate is small.
citation-role summary
citation-polarity summary
years
2026 6roles
background 1polarities
background 1representative citing papers
A Langevin training trajectory's chance of occupying a small failure region relaxes to about twice its tiny stationary value after a burn-in of order the dimension, unless the region's geometry gives a faster local relaxation rate.
Derives second-order path-kernel interpolation formulas for gradient descent, SGD, and momentum training, adding curvature terms and a concentration estimate around the expected prediction.
Four characterizations of irreversibility in training algorithms are equivalent to leading order in step size and produce an emergent force that breaks reparametrization symmetries while favoring minimum entropy production trajectories.
The paper defines computational effort as the number of gradient descent steps to reach target accuracy with high probability, shows large learning rates minimize this effort across models, and identifies phase transitions in optimal training strategies.
citing papers explorer
-
Estimating Implicit Regularization in Deep Learning
Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
-
Avoiding unsafe sets when training with Langevin Dynamics
A Langevin training trajectory's chance of occupying a small failure region relaxes to about twice its tiny stationary value after a burn-in of order the dimension, unless the region's geometry gives a faster local relaxation rate.
-
Second-Order Path Kernel Interpolation Formulas in Machine Learning
Derives second-order path-kernel interpolation formulas for gradient descent, SGD, and momentum training, adding curvature terms and a concentration estimate around the expected prediction.
-
Thermodynamic Irreversibility of Training Algorithms
Four characterizations of irreversibility in training algorithms are equivalent to leading order in step size and produce an emergent force that breaks reparametrization symmetries while favoring minimum entropy production trajectories.
-
Gradient-Descent Steps to Success over Mean Accuracy: A Paradigm Shift for ML
The paper defines computational effort as the number of gradient descent steps to reach target accuracy with high probability, shows large learning rates minimize this effort across models, and identifies phase transitions in optimal training strategies.
- Convergence of difference inclusions: a diameter criterion and step-size conditions