DP-NGD enables second-order optimization under differential privacy by decoupling curvature estimation onto public data, performing isotropic DP operations in a whitened space, and dynamically clamping curvature eigenvalues to prevent instability.
Adam with model exponential moving average is effective for nonconvex optimization
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this work, we offer a theoretical analysis of two modern optimization techniques for training large and complex models: (i) adaptive optimization algorithms, such as Adam, and (ii) the model exponential moving average (EMA). Specifically, we demonstrate that a clipped version of Adam with model EMA achieves the optimal convergence rates in various nonconvex optimization settings, both smooth and nonsmooth. Moreover, when the scale varies significantly across different coordinates, we demonstrate that the coordinate-wise adaptivity of Adam is provably advantageous. Notably, unlike previous analyses of Adam, our analysis crucially relies on its core elements -- momentum and discounting factors -- as well as model EMA, motivating their wide applications in practice.
years
2026 2representative citing papers
Establishes a central limit theorem for averaged Adam with n^{-1/2} convergence rate to an attracting zero and covariance determined by the algorithm at the attractor.
citing papers explorer
-
Differentially Private Natural Gradient Descent
DP-NGD enables second-order optimization under differential privacy by decoupling curvature estimation onto public data, performing isotropic DP operations in a whitened space, and dynamically clamping curvature eigenvalues to prevent instability.
-
Central limit theorem for the averaged Adam optimizer
Establishes a central limit theorem for averaged Adam with n^{-1/2} convergence rate to an attracting zero and covariance determined by the algorithm at the attractor.