Pith. sign in

REVIEW 1 cited by

The Limit of the Batch Size

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.08517 v1 pith:GDMKE6GD submitted 2020-06-15 cs.LG cs.CVcs.DCstat.ML

classification cs.LGcs.CVcs.DCstat.ML
keywords batchoptimizationsizetrainingdetaileddiffusionhofferimagenet
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-batch training is an efficient approach for current distributed deep learning systems. It has enabled researchers to reduce the ImageNet/ResNet-50 training from 29 hours to around 1 minute. In this paper, we focus on studying the limit of the batch size. We think it may provide a guidance to AI supercomputer and algorithm designers. We provide detailed numerical optimization instructions for step-by-step comparison. Moreover, it is important to understand the generalization and optimization performance of huge batch training. Hoffer et al. introduced "ultra-slow diffusion" theory to large-batch training. However, our experiments show contradictory results with the conclusion of Hoffer et al. We provide comprehensive experimental results and detailed analysis to study the limitations of batch size scaling and "ultra-slow diffusion" theory. For the first time we scale the batch size on ImageNet to at least a magnitude larger than all previous work, and provide detailed studies on the performance of many state-of-the-art optimization schemes under this setting. We propose an optimization recipe that is able to improve the top-1 test accuracy by 18% compared to the baseline.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Collaborative Batch Size Optimization for Federated Learning

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A distributed randomized binary search lets federated learning clients find a hardware-safe shared batch size within a few rounds, speeding up training compared to a small default batch size.

Pith tools