Online Batch Selection for Faster Training of Neural Networks

arxiv: 1511.06343 · v4 · pith:XZJUB4DSnew · submitted 2015-11-19 · 💻 cs.LG · cs.NE· math.OC

Online Batch Selection for Faster Training of Neural Networks

Ilya Loshchilov , Frank Hutter This is my paper

classification 💻 cs.LG cs.NEmath.OC

keywords batchlossselectionbatchesdatapointsdatasetonlineadadelta

0 comments p. Extension

pith:XZJUB4DS Add to your LaTeX paper

What is a Pith Number?

\usepackage{pith}
\pithnumber{XZJUB4DS}

Prints a linked pith:XZJUB4DS badge after your title and writes the identifier into PDF metadata. Compiles on arXiv with no extra files. Learn more

read the original abstract

Deep neural networks are commonly trained using stochastic non-convex optimization procedures, which are driven by gradient information estimated on fractions (batches) of the dataset. While it is commonly accepted that batch size is an important parameter for offline tuning, the benefits of online selection of batches remain poorly understood. We investigate online batch selection strategies for two state-of-the-art methods of stochastic gradient-based optimization, AdaDelta and Adam. As the loss function to be minimized for the whole dataset is an aggregation of loss functions of individual datapoints, intuitively, datapoints with the greatest loss should be considered (selected in a batch) more frequently. However, the limitations of this intuition and the proper control of the selection pressure over time are open questions. We propose a simple strategy where all datapoints are ranked w.r.t. their latest known loss value and the probability to be selected decays exponentially as a function of rank. Our experimental results on the MNIST dataset suggest that selecting batches speeds up both AdaDelta and Adam by a factor of about 5.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Disagreement-Regularized Importance Sampling for Adversarial Label Corruption
cs.LG 2026-05 unverdicted novelty 7.0

DR-IS selects low-contamination subsets via bounded rank-disagreement in proxy ensembles under an ε-contamination model, with O(√(log(N/δ)/K)) concentration rates that certify separation when the expectation gap Δ' is...
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
cs.LG 2026-05 unverdicted novelty 6.0

Dr. Post-Training reframes general data as a data-induced regularizer for LLM post-training updates, yielding a family of methods that outperform data-selection baselines on SFT, RLHF, and RLVR tasks.
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training
cs.LG 2026-04 conditional novelty 6.0

Data Warmup accelerates diffusion training on ImageNet by scheduling images from low to high complexity via a foreground-based metric and temperature-controlled sampler, improving FID and IS scores faster than uniform...
Variance Matters: Improving Domain Adaptation via Stratified Sampling
cs.LG 2025-12 unverdicted novelty 6.0

VaRDASS improves unsupervised domain adaptation by using stratified sampling to reduce variance in discrepancy estimation for measures like correlation alignment and MMD, with derived error bounds, an optimality proof...
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
cs.CV 2025-02 unverdicted novelty 4.0

Step-Video-T2V describes a 30B-parameter text-to-video model with custom Video-VAE, 3D DiT, flow matching, and Video-DPO that claims state-of-the-art results on a new internal benchmark.