For convex stochastic optimization under (L0,L1)-smoothness, clipped and normalized SGD (and zero-order variants) obtain linear-rate terms in their convergence bounds, and in the L0=0 regime NSGD achieves logarithmic iteration complexity at the cost of large batches.
Neural optimizer search with reinforcement learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Power of Generalized Smoothness in Stochastic Convex Optimization: First- and Zero-Order Algorithms
For convex stochastic optimization under (L0,L1)-smoothness, clipped and normalized SGD (and zero-order variants) obtain linear-rate terms in their convergence bounds, and in the L0=0 regime NSGD achieves logarithmic iteration complexity at the cost of large batches.