The paper introduces (L,\bar L)-anisotropic smoothness and proves O(1/K) convergence rates for nonlinearly preconditioned gradient methods, unifying gradient clipping, Adam, and Adagrad under one theory.
Thus we can further bound (30): AK ξ(xK)−ξ(x ⋆) ≤ D0 K−1X k=0 a2 k+1 Ak+1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Nonlinearly Preconditioned Gradient Methods under Generalized Smoothness
The paper introduces (L,\bar L)-anisotropic smoothness and proves O(1/K) convergence rates for nonlinearly preconditioned gradient methods, unifying gradient clipping, Adam, and Adagrad under one theory.