Gradient descent with the new adaptive step size reaches near-optimal or first-known convergence rates for ℓ-smooth functions, including the previously open quadratic-growth case.
This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
math.OC 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Toward a Unified Theory of Gradient Descent under Generalized Smoothness
Gradient descent with the new adaptive step size reaches near-optimal or first-known convergence rates for ℓ-smooth functions, including the previously open quadratic-growth case.