REVIEW 2 cited by
Stochastic Training is Not Necessary for Generalization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
It is widely believed that the implicit regularization of SGD is fundamental to the impressive generalization behavior we observe in neural networks. In this work, we demonstrate that non-stochastic full-batch training can achieve comparably strong performance to SGD on CIFAR-10 using modern architectures. To this end, we show that the implicit regularization of SGD can be completely replaced with explicit regularization even when comparing against a strong and well-researched baseline. Our observations indicate that the perceived difficulty of full-batch training may be the result of its optimization properties and the disproportionate time and effort spent by the ML community tuning optimizers and hyperparameters for small-batch training.
Forward citations
Cited by 2 Pith papers
-
Parameter Symmetry Potentially Unifies Deep Learning Theory
This position paper argues that parameter symmetry breaking and restoration unify three hierarchies in deep learning: learning dynamics, model complexity, and representation formation.
-
Knowledge distillation as a pathway toward next-generation intelligent ecohydrological modeling systems
The paper proposes a three-phase knowledge distillation pathway (behavioral, structural, cognitive) for embedding process-based ecohydrological models into AI, with limited Samish watershed demonstrations.
Discussion (0). Continue with ORCID to comment.