Pith. sign in

REVIEW

Distributed Momentum for Byzantine-resilient Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.00010 v2 pith:KPJNROEI submitted 2020-02-28 cs.LG cs.CRcs.DC

classification cs.LGcs.CRcs.DC
keywords momentumaggregationdistributedrobustnessserverbenefitsgradientlinear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Momentum is a variant of gradient descent that has been proposed for its benefits on convergence. In a distributed setting, momentum can be implemented either at the server or the worker side. When the aggregation rule used by the server is linear, commutativity with addition makes both deployments equivalent. Robustness and privacy are however among motivations to abandon linear aggregation rules. In this work, we demonstrate the benefits on robustness of using momentum at the worker side. We first prove that computing momentum at the workers reduces the variance-norm ratio of the gradient estimation at the server, strengthening Byzantine resilient aggregation rules. We then provide an extensive experimental demonstration of the robustness effect of worker-side momentum on distributed SGD.

Discussion (0). Continue with ORCID to comment.

Pith tools