The accelerated Allen-Cahn equation formally converges to the hyperbolic interface law ∂_t v = (1-v^2)(h-αv), and a large-step FISTA discretization empirically accelerates Ginzburg-Landau minimization.
Nesterov acceleration in benignly non-convex landscapes
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
While momentum-based optimization algorithms are commonly used in the notoriously non-convex optimization problems of deep learning, their analysis has historically been restricted to the convex and strongly convex setting. In this article, we partially close this gap between theory and practice and demonstrate that virtually identical guarantees can be obtained in optimization problems with a `benign' non-convexity. We show that these weaker geometric assumptions are well justified in overparametrized deep learning, at least locally. Variations of this result are obtained for a continuous time model of Nesterov's accelerated gradient descent algorithm (NAG), the classical discrete time version of NAG, and versions of NAG with stochastic gradient estimates with purely additive noise and with noise that exhibits both additive and multiplicative scaling.
fields
math.AP 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Momentum-based minimization of the Ginzburg-Landau functional on Euclidean spaces and graphs
The accelerated Allen-Cahn equation formally converges to the hyperbolic interface law ∂_t v = (1-v^2)(h-αv), and a large-step FISTA discretization empirically accelerates Ginzburg-Landau minimization.