Pith. sign in

REVIEW 1 cited by

Asynchronous Stochastic Gradient Descent with Delay Compensation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1609.08326 v6 pith:F54PVYPA submitted 2016-09-27 cs.LG cs.DC

classification cs.LGcs.DC
keywords gradientasgdasynchronousdelayalgorithmdc-asgddelayeddescent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the fast development of deep learning, it has become common to learn big neural networks using massive training data. Asynchronous Stochastic Gradient Descent (ASGD) is widely adopted to fulfill this task for its efficiency, which is, however, known to suffer from the problem of delayed gradients. That is, when a local worker adds its gradient to the global model, the global model may have been updated by other workers and this gradient becomes "delayed". We propose a novel technology to compensate this delay, so as to make the optimization behavior of ASGD closer to that of sequential SGD. This is achieved by leveraging Taylor expansion of the gradient function and efficient approximation to the Hessian matrix of the loss function. We call the new algorithm Delay Compensated ASGD (DC-ASGD). We evaluated the proposed algorithm on CIFAR-10 and ImageNet datasets, and the experimental results demonstrate that DC-ASGD outperforms both synchronous SGD and asynchronous SGD, and nearly approaches the performance of sequential SGD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decentralized Federated Learning by Partial Message Exchange

    cs.LG 2026-03 reject novelty 4.0 of 10

    PaME combines random coordinate exchange with a growing-penalty schedule, claiming linear convergence under two mild assumptions, but its key parameter condition is never satisfied by its own experiments and the limit...

Pith tools