Gradient descent with the new adaptive step size reaches near-optimal or first-known convergence rates for ℓ-smooth functions, including the previously open quadratic-growth case.
For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Toward a Unified Theory of Gradient Descent under Generalized Smoothness
Gradient descent with the new adaptive step size reaches near-optimal or first-known convergence rates for ℓ-smooth functions, including the previously open quadratic-growth case.