Pith. sign in

REVIEW 13 cited by

Measuring the Effects of Data Parallelism on Neural Network Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.03600 v3 pith:V6CNNA62 submitted 2018-11-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords trainingbatchdataneuralnetworksizeavailableeffects
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent hardware developments have dramatically increased the scale of data parallelism available for neural network training. Among the simplest ways to harness next-generation hardware is to increase the batch size in standard mini-batch neural network training algorithms. In this work, we aim to experimentally characterize the effects of increasing the batch size on training time, as measured by the number of steps necessary to reach a goal out-of-sample error. We study how this relationship varies with the training algorithm, model, and data set, and find extremely large variation between workloads. Along the way, we show that disagreements in the literature on how batch size affects model quality can largely be explained by differences in metaparameter tuning and compute budgets at different batch sizes. We find no evidence that larger batch sizes degrade out-of-sample performance. Finally, we discuss the implications of our results on efforts to train neural networks much faster in the future. Our experimental data is publicly available as a database of 71,638,836 loss measurements taken over the course of training for 168,160 individual models across 35 workloads.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    The authors derive a Maximally Scale-Stable Parameterization (MSSP) for MoE models that achieves robust learning-rate transfer and monotonic performance gains with scale across co-scaling regimes of width, experts, an...

  2. FinBERT: Financial Sentiment Analysis with Pre-trained Language Models

    cs.CL 2019-08 unverdicted novelty 7.0 of 10

    FinBERT adapts BERT to the financial domain and outperforms prior state-of-the-art methods on financial sentiment analysis tasks.

  3. The Fourth Quadrant: A Stylized View of Benign Misfitting

    cs.LG 2026-08 conditional novelty 6.0 of 10

    In a stylized single-spike linear model, useful span predictors in the window d/gamma^2 << n << d/gamma are forced to overshoot the training labels, so good test error comes together with large training error.

  4. MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training

    cs.LG 2025-08 conditional novelty 6.0 of 10

    MERIT, a max-norm and element-wise trust-ratio optimizer, improves large-batch GPT-2 and Llama training and matches small-batch downstream scores at 6k batch size.

  5. Scaling Laws for Transfer

    cs.LG 2021-02 unverdicted novelty 6.0 of 10

    Effective data transferred from pre-training to fine-tuning is described by a power law in model parameter count and fine-tuning dataset size, acting like a multiplier on the fine-tuning data.

  6. Hierarchical Federated Learning Across Heterogeneous Cellular Networks

    cs.LG 2019-09 conditional novelty 6.0 of 10

    A hierarchical federated learning framework with gradient sparsification reduces modeled communication latency in heterogeneous cellular networks while keeping CIFAR-10 accuracy close to a flat baseline.

  7. Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

    cs.LG 2019-04 conditional novelty 6.0 of 10

    LAMB optimizer trains BERT with batch size 32868, reducing training time to 76 minutes on TPUv3 Pod without performance loss.

  8. Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification

    cs.LG 2019-09 unverdicted novelty 5.0 of 10

    Non-identical data distributions degrade federated averaging accuracy on visual classification, but server momentum raises CIFAR-10 accuracy from 30.1% to 76.9% in the most skewed regimes.

  9. Mix & Match: training convnets with mixed image sizes for improved accuracy, speed and scale resiliency

    cs.CV 2019-08 conditional novelty 5.0 of 10

    MixSize training makes ImageNet classifiers resilient to smaller test images, matching baseline top-1 accuracy at 160x160 with about half the inference compute, while optionally improving accuracy or training speed.

  10. Fast Training of Sparse Graph Neural Networks on Dense Hardware

    stat.ML 2019-06 unverdicted novelty 5.0 of 10

    Techniques enable training the sparse GNN from Allamanis et al. [2018] on dense TPU hardware in 13 minutes versus a full day originally.

  11. What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.

  12. Trustworthy Efficient Communication for Distributed Learning using LQ-SGD Algorithm

    cs.LG 2025-06 conditional novelty 4.0 of 10

    LQ-SGD combines low-rank gradient compression with log quantization, claiming roughly 4x lower communication than PowerSGD while preserving accuracy and improving gradient-inversion resistance.

  13. SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Higher learning-rate-to-batch-size ratios improve reasoning-task performance in small language models, and the resulting SmolTulu-1.7B reports top sub-2B scores on IFEval and GSM8K.

Pith tools