Pith. sign in

REVIEW 18 cited by

Why gradient clipping accelerates training: A theoretical justification for adaptivity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.11881 v2 pith:ENLPZKPT submitted 2019-05-28 math.OC cs.LG

Why gradient clipping accelerates training: A theoretical justification for adaptivity

classification math.OC cs.LG
keywords gradientsmoothnesstrainingneuralclippingtheoreticalalgorithmscondition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We provide a theoretical explanation for the effectiveness of gradient clipping in training deep neural networks. The key ingredient is a new smoothness condition derived from practical neural network training examples. We observe that gradient smoothness, a concept central to the analysis of first-order optimization algorithms that is often assumed to be a constant, demonstrates significant variability along the training trajectory of deep neural networks. Further, this smoothness positively correlates with the gradient norm, and contrary to standard assumptions in the literature, it can grow with the norm of the gradient. These empirical observations limit the applicability of existing theoretical analyses of algorithms that rely on a fixed bound on smoothness. These observations motivate us to introduce a novel relaxation of gradient smoothness that is weaker than the commonly used Lipschitz smoothness assumption. Under the new condition, we prove that two popular methods, namely, \emph{gradient clipping} and \emph{normalized gradient}, converge arbitrarily faster than gradient descent with fixed stepsize. We further explain why such adaptively scaled gradient methods can accelerate empirical convergence and verify our results empirically in popular neural network training settings.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Stochastic Non-Smooth Convex Optimization with Unbounded Gradients

    math.OC 2026-05 unverdicted novelty 8.0

    Introduces generalized Lipschitz class and shows clipped AdamW outperforms SGD and AdaGrad for stochastic convex optimization under this and related assumptions.

  2. Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime

    cs.LG 2026-06 unverdicted novelty 7.0

    Proves GD convergence to stationary point neighborhoods for general NN architectures beyond NTK via block-level analysis, analyticity, and local smoothness conditions.

  3. OptMuon: Closed-Loop Orthogonalized Momentum Methods for Stochastic Optimization with Zero-Noise Optimality

    math.OC 2026-06 unverdicted novelty 7.0

    OptMuon combines orthogonalized momentum with closed-loop adaptation to achieve noise-adaptive convergence rates that automatically become near-optimal deterministic first-order rates without retuning when noise vanishes.

  4. Stochastic Non-Smooth Convex Optimization with Unbounded Gradients

    math.OC 2026-05 unverdicted novelty 7.0

    Clipped AdamW with exponentially weighted accumulation achieves superior global convergence rates for convex stochastic generalized Lipschitz optimization compared to SGD and AdaGrad.

  5. Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum-Gladyshev Noise

    cs.LG 2026-05 unverdicted novelty 7.0

    Normalized momentum SGD and variance-reduced STORM achieve O(ε^{-6}) and O(ε^{-4}) oracle complexities respectively under quadratic distance-dependent noise in nonconvex stochastic optimization.

  6. Newton methods beyond Hessian Lipschitz continuity: A nonlinear preconditioning approach

    math.OC 2026-05 unverdicted novelty 7.0

    Nonlinear preconditioning extends Newton methods to objectives lacking Hessian Lipschitz continuity by analyzing a transformed mapping under a relaxed smoothness condition, with superlinear convergence and O(ε^{-3/2})...

  7. The Multi-Block DC Function Class: Theory, Algorithms, and Applications

    math.OC 2026-04 unverdicted novelty 7.0

    The Multi-Block DC class admits polynomial-size DC decompositions for problems that require exponential size under standard DC programming and supplies explicit constructive formulations for deep ReLU networks togethe...

  8. Distribution-Aware Robust Bilevel Optimization: Quantile-Guided Huber Updates in Two-Timescale Stochastic Approximation

    cs.LG 2026-06 unverdicted novelty 6.0

    RQ-TTSA achieves O(T^{-(p-1)/(3p-2)}) convergence for nonconvex-strongly convex bilevel optimization under heavy-tailed noise (p in (1,2]) via quantile-guided Huber clipping and shows empirical gains on vision, games,...

  9. Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case

    math.OC 2026-06 unverdicted novelty 6.0

    Convergence rates are derived for Muon-type methods with inexact LMO in the degenerate case under novel assumptions and layer-wise (L^0, L^1)-smoothness for non-convex and star-convex objectives with weight decay.

  10. OptMuon: Closed-Loop Orthogonalized Momentum Methods for Stochastic Optimization with Zero-Noise Optimality

    math.OC 2026-06 unverdicted novelty 6.0

    OptMuon combines orthogonalized momentum with trajectory-dependent AdaGrad-Norm adaptation to obtain expected-stationarity rates of order T^{-1/2} + sigma^{1/2}T^{-1/4} or T^{-1/2} + sigma^{1/3}T^{-1/3} that reduce to...

  11. Distributionally Robust Multi-Objective Optimization

    cs.LG 2026-05 unverdicted novelty 6.0

    DR-MOO adds distributional robustness to multi-objective optimization and gives single-loop MGDA algorithms reaching epsilon-Pareto-stationary points in O(epsilon^{-4}) samples for nonconvex problems.

  12. Cost-Aware Learning

    cs.LG 2026-04 unverdicted novelty 6.0

    Cost-aware SGD achieves target error with lower total sampling cost than standard methods, and Cost-Aware GRPO reduces token usage by up to 30% in LLM reinforcement learning while matching baseline performance.

  13. Adaptive Federated Optimization

    cs.LG 2020-02 unverdicted novelty 6.0

    Proposes federated adaptive optimizers (FedAdagrad, FedAdam, FedYogi) with convergence analysis for non-convex objectives under data heterogeneity and reports empirical gains over FedAvg.

  14. Revisiting Privacy Amplification by Subsampling in Selective Release DPSGD

    cs.LG 2026-06 unverdicted novelty 5.0

    DPSR-CG corrects the privacy accounting for selective release in DPSGD by addressing sampling probability variation and reports strong empirical results on MNIST, CIFAR-10, IMDB, and FMNIST while claiming strict privacy.

  15. Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives

    math.OC 2026-05 unverdicted novelty 5.0

    Proximal stochastic spectral preconditioning converges for nonconvex constrained objectives under heavy-tailed noise, with a variance-reduced version achieving faster rates and a refined analysis of Muon iterations.

  16. Cost-Aware Learning

    cs.LG 2026-04 unverdicted novelty 5.0

    Cost-Aware SGD samples by gradient-norm-to-cost ratio and is instantiated as Cost-Aware GRPO for length-dependent policy gradients, reducing tokens used in LLM RL while matching baseline accuracy.

  17. Frank-Wolfe Algorithms for (L0, L1)-smooth functions

    math.OC 2025-10 unverdicted novelty 5.0

    Proposes (L0, L1)-Frank-Wolfe and adaptive variant claiming superior convergence rates for (L0, L1)-smooth objectives over classical Frank-Wolfe.

  18. Frank-Wolfe Algorithms for (L0, L1)-smooth functions

    math.OC 2025-10 unverdicted novelty 5.0

    A new (L0, L1)-Frank-Wolfe algorithm and its adaptive version are proposed for (L0, L1)-smooth optimization, with claims of better theoretical convergence rates and practical advantages over standard Frank-Wolfe methods.