REVIEW 2 cited by
Augment your batch: better training with larger batches
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large-batch SGD is important for scaling training of deep neural networks. However, without fine-tuning hyperparameter schedules, the generalization of the model may be hampered. We propose to use batch augmentation: replicating instances of samples within the same batch with different data augmentations. Batch augmentation acts as a regularizer and an accelerator, increasing both generalization and performance scaling. We analyze the effect of batch augmentation on gradient variance and show that it empirically improves convergence for a wide variety of deep neural networks and datasets. Our results show that batch augmentation reduces the number of necessary SGD updates to achieve the same accuracy as the state-of-the-art. Overall, this simple yet effective method enables faster training and better generalization by allowing more computational resources to be used concurrently.
Forward citations
Cited by 2 Pith papers
-
Galileo: Learning Global & Local Features of Many Remote Sensing Modalities
A single multimodal transformer, Galileo, jointly learns global and local features from optical, radar, elevation, weather, and land-cover inputs and outperforms specialized models on eleven benchmarks.
-
Mix & Match: training convnets with mixed image sizes for improved accuracy, speed and scale resiliency
MixSize training makes ImageNet classifiers resilient to smaller test images, matching baseline top-1 accuracy at 160x160 with about half the inference compute, while optionally improving accuracy or training speed.
Discussion (0). Continue with ORCID to comment.