Gradient descent with the new adaptive step size reaches near-optimal or first-known convergence rates for ℓ-smooth functions, including the previously open quadratic-growth case.
Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Toward a Unified Theory of Gradient Descent under Generalized Smoothness
Gradient descent with the new adaptive step size reaches near-optimal or first-known convergence rates for ℓ-smooth functions, including the previously open quadratic-growth case.