Distributed momentum SGD achieves last-iterate almost sure and mean-square convergence of the gradient norm in non-convex settings under Robbins-Monro step sizes.
Distributed training strategies for the structured perceptron,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Convergence Analysis of the Last Iterate in Distributed Stochastic Gradient Descent with Momentum
Distributed momentum SGD achieves last-iterate almost sure and mean-square convergence of the gradient norm in non-convex settings under Robbins-Monro step sizes.