Pith. sign in

REVIEW 11 cited by

Implicit Gradient Regularization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.11162 v3 pith:AQH7M24E submitted 2020-09-23 cs.LG stat.ML

classification cs.LGstat.ML
keywords gradientregularizationdescentimplicitanalysisbackwarderrorexplicit
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Gradient descent can be surprisingly good at optimizing deep neural networks without overfitting and without explicit regularization. We find that the discrete steps of gradient descent implicitly regularize models by penalizing gradient descent trajectories that have large loss gradients. We call this Implicit Gradient Regularization (IGR) and we use backward error analysis to calculate the size of this regularization. We confirm empirically that implicit gradient regularization biases gradient descent toward flat minima, where test errors are small and solutions are robust to noisy parameter perturbations. Furthermore, we demonstrate that the implicit gradient regularization term can be used as an explicit regularizer, allowing us to control this gradient regularization directly. More broadly, our work indicates that backward error analysis is a useful theoretical approach to the perennial question of how learning rate, model size, and parameter regularization interact to determine the properties of overparameterized models optimized with gradient descent.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Universal Convergence of Backward Error in Linear System Solvers

    math.NA 2026-04 unverdicted novelty 8.0 of 10

    Richardson iteration achieves universal 1/k backward error on PSD systems, enabling O(n²/ε) solvers; MINBERR reaches O(1/k²) rate and O(n²/√ε) complexity, with empirical O(1/k) extension to general systems.

  2. Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Long-context training shifts language models from parametric knowledge to context reliance, producing an inverted-U in pretraining performance and context addiction in supervised fine-tuning.

  3. Avoiding unsafe sets when training with Langevin Dynamics

    cs.LG 2026-07 accept novelty 6.0 of 10

    Langevin training trajectories on strongly convex losses avoid geometrically isolated failure regions with probability exponentially small in dimension after an O(d) burn-in, with a local spectral rate controlling tra...

  4. Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Regularizing the policy-gradient norm during RLHF/RLVR training biases the policy toward flat optima where the proxy reward stays accurate, mitigating reward hacking better than a KL penalty.

  5. Quantitative Understanding of PDF Fits and their Uncertainties

    hep-ph 2025-12 conditional novelty 6.0 of 10

    After an initial transient, a PDF-fitting neural network's output obeys f_t = U(t) f_0 + V(t) Y, a linear blend of the initial network and the data with explicit time-dependent operators.

  6. What Can Grokking Teach Us About Learning Under Nonstationarity?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Periodically increasing the effective learning rate while constraining parameter norms induces feature-learning dynamics and mitigates primacy bias in grokking, warm-starting, and reinforcement learning.

  7. Neural Thermodynamic Laws for Large Language Model Training

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Under a river-valley model of the loss landscape, the paper shows that valley fluctuations behave like heat, with learning rate as temperature, and derives a 1/t optimal decay schedule.

  8. Explicit Eigenvalue Regularization Improves Sharpness-Aware Minimization

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Eigen-SAM improves Sharpness-Aware Minimization by explicitly aligning the perturbation with the top Hessian eigenvector, supported by a third-order SDE analysis and consistent small accuracy gains on CIFAR, SVHN, and...

  9. Optimizers Qualitatively Alter Solutions And We Should Leverage This

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Deep learning optimizers should be designed to induce desired solution properties, not just convergence speed; different optimizers demonstrably land in qualitatively different minima.

  10. Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability

    math.NA 2025-11 conditional novelty 3.0 of 10

    The late-stage noise floor of preconditioned SGD is the product of the M-metric condition number and the preconditioned noise level, so the design goal is to improve conditioning while dampening noise.

  11. CGD: Modifying the Loss Landscape by Gradient Regularization

    math.OC 2025-04 conditional novelty 3.0 of 10

    CGD, gradient descent on a gradient-norm-penalized objective, has a proven linear convergence rate and practical finite-difference and quasi-Newton variants, though the core idea matches explicit gradient regularization.

Pith tools