Richardson and a new Krylov method MINBERR achieve universal (condition-free) backward-error rates 1/k and O(1/k^{2}) for PSD linear systems, with a near-universal O(log n / k) extension to general systems.
Implicit gradient regularization
8 Pith papers cite this work. Polarity classification is still indexing.
abstract
Gradient descent can be surprisingly good at optimizing deep neural networks without overfitting and without explicit regularization. We find that the discrete steps of gradient descent implicitly regularize models by penalizing gradient descent trajectories that have large loss gradients. We call this Implicit Gradient Regularization (IGR) and we use backward error analysis to calculate the size of this regularization. We confirm empirically that implicit gradient regularization biases gradient descent toward flat minima, where test errors are small and solutions are robust to noisy parameter perturbations. Furthermore, we demonstrate that the implicit gradient regularization term can be used as an explicit regularizer, allowing us to control this gradient regularization directly. More broadly, our work indicates that backward error analysis is a useful theoretical approach to the perennial question of how learning rate, model size, and parameter regularization interact to determine the properties of overparameterized models optimized with gradient descent.
citation-role summary
citation-polarity summary
years
2026 8roles
background 2polarities
background 2representative citing papers
Online kernel regression equals offline regression with shifted targets; correcting the targets lets online learning match offline performance and outperform true targets in continual image classification.
Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
A Langevin training trajectory's chance of occupying a small failure region relaxes to about twice its tiny stationary value after a burn-in of order the dimension, unless the region's geometry gives a faster local relaxation rate.
Derives second-order path-kernel interpolation formulas for gradient descent, SGD, and momentum training, adding curvature terms and a concentration estimate around the expected prediction.
Four characterizations of irreversibility in training algorithms are equivalent to leading order in step size and produce an emergent force that breaks reparametrization symmetries while favoring minimum entropy production trajectories.
Derives optimality constraints for nonnegative joint dictionary learning that explain observed SAE behaviors such as feature splitting, absorption, and dense antipodal features.
Lectures reviewing three established numerical methods for inverse problems in extracting PDFs and spectral functions from lattice QCD and experimental data.
citing papers explorer
-
Towards Universal Convergence of Backward Error in Linear System Solvers
Richardson and a new Krylov method MINBERR achieve universal (condition-free) backward-error rates 1/k and O(1/k^{2}) for PSD linear systems, with a near-universal O(log n / k) extension to general systems.
-
Characterizing and Correcting Effective Target Shift in Online Learning
Online kernel regression equals offline regression with shifted targets; correcting the targets lets online learning match offline performance and outperform true targets in continual image classification.
-
Estimating Implicit Regularization in Deep Learning
Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
-
Avoiding unsafe sets when training with Langevin Dynamics
A Langevin training trajectory's chance of occupying a small failure region relaxes to about twice its tiny stationary value after a burn-in of order the dimension, unless the region's geometry gives a faster local relaxation rate.
-
Second-Order Path Kernel Interpolation Formulas in Machine Learning
Derives second-order path-kernel interpolation formulas for gradient descent, SGD, and momentum training, adding curvature terms and a concentration estimate around the expected prediction.
-
Thermodynamic Irreversibility of Training Algorithms
Four characterizations of irreversibility in training algorithms are equivalent to leading order in step size and produce an emergent force that breaks reparametrization symmetries while favoring minimum entropy production trajectories.
-
How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations
Derives optimality constraints for nonnegative joint dictionary learning that explain observed SAE behaviors such as feature splitting, absorption, and dense antipodal features.
-
Some Inverse Problems in Particle Physics
Lectures reviewing three established numerical methods for inverse problems in extracting PDFs and spectral functions from lattice QCD and experimental data.