REVIEW 13 cited by
Measuring the Effects of Data Parallelism on Neural Network Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent hardware developments have dramatically increased the scale of data parallelism available for neural network training. Among the simplest ways to harness next-generation hardware is to increase the batch size in standard mini-batch neural network training algorithms. In this work, we aim to experimentally characterize the effects of increasing the batch size on training time, as measured by the number of steps necessary to reach a goal out-of-sample error. We study how this relationship varies with the training algorithm, model, and data set, and find extremely large variation between workloads. Along the way, we show that disagreements in the literature on how batch size affects model quality can largely be explained by differences in metaparameter tuning and compute budgets at different batch sizes. We find no evidence that larger batch sizes degrade out-of-sample performance. Finally, we discuss the implications of our results on efforts to train neural networks much faster in the future. Our experimental data is publicly available as a database of 71,638,836 loss measurements taken over the course of training for 168,160 individual models across 35 workloads.
Forward citations
Cited by 13 Pith papers
-
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
The authors derive a Maximally Scale-Stable Parameterization (MSSP) for MoE models that achieves robust learning-rate transfer and monotonic performance gains with scale across co-scaling regimes of width, experts, an...
-
FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
FinBERT adapts BERT to the financial domain and outperforms prior state-of-the-art methods on financial sentiment analysis tasks.
-
The Fourth Quadrant: A Stylized View of Benign Misfitting
In a stylized single-spike linear model, useful span predictors in the window d/gamma^2 << n << d/gamma are forced to overshoot the training labels, so good test error comes together with large training error.
-
MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
MERIT, a max-norm and element-wise trust-ratio optimizer, improves large-batch GPT-2 and Llama training and matches small-batch downstream scores at 6k batch size.
-
Scaling Laws for Transfer
Effective data transferred from pre-training to fine-tuning is described by a power law in model parameter count and fine-tuning dataset size, acting like a multiplier on the fine-tuning data.
-
Hierarchical Federated Learning Across Heterogeneous Cellular Networks
A hierarchical federated learning framework with gradient sparsification reduces modeled communication latency in heterogeneous cellular networks while keeping CIFAR-10 accuracy close to a flat baseline.
-
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
LAMB optimizer trains BERT with batch size 32868, reducing training time to 76 minutes on TPUv3 Pod without performance loss.
-
Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification
Non-identical data distributions degrade federated averaging accuracy on visual classification, but server momentum raises CIFAR-10 accuracy from 30.1% to 76.9% in the most skewed regimes.
-
Mix & Match: training convnets with mixed image sizes for improved accuracy, speed and scale resiliency
MixSize training makes ImageNet classifiers resilient to smaller test images, matching baseline top-1 accuracy at 160x160 with about half the inference compute, while optionally improving accuracy or training speed.
-
Fast Training of Sparse Graph Neural Networks on Dense Hardware
Techniques enable training the sparse GNN from Allamanis et al. [2018] on dense TPU hardware in 13 minutes versus a full day originally.
-
What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness
Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.
-
Trustworthy Efficient Communication for Distributed Learning using LQ-SGD Algorithm
LQ-SGD combines low-rank gradient compression with log quantization, claiming roughly 4x lower communication than PowerSGD while preserving accuracy and improving gradient-inversion resistance.
-
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
Higher learning-rate-to-batch-size ratios improve reasoning-task performance in small language models, and the resulting SmolTulu-1.7B reports top sub-2B scores on IFEval and GSM8K.
Discussion (0). Continue with ORCID to comment.