REVIEW 1 cited by
GPU Asynchronous Stochastic Gradient Descent to Speed Up Neural Network Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability to train large-scale neural networks has resulted in state-of-the-art performance in many areas of computer vision. These results have largely come from computational break throughs of two forms: model parallelism, e.g. GPU accelerated training, which has seen quick adoption in computer vision circles, and data parallelism, e.g. A-SGD, whose large scale has been used mostly in industry. We report early experiments with a system that makes use of both model parallelism and data parallelism, we call GPU A-SGD. We show using GPU A-SGD it is possible to speed up training of large convolutional neural networks useful for computer vision. We believe GPU A-SGD will make it possible to train larger networks on larger training sets in a reasonable amount of time.
Forward citations
Cited by 1 Pith paper
-
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
Fault-aware scaling, tile reordering, and fine-tuning restore near-original accuracy for many single stuck-at-bit faults in systolic-array neural network accelerators, in simulation.
Discussion (0). Continue with ORCID to comment.