REVIEW 3 cited by
A Study of Gradient Variance in Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The impact of gradient noise on training deep models is widely acknowledged but not well understood. In this context, we study the distribution of gradients during training. We introduce a method, Gradient Clustering, to minimize the variance of average mini-batch gradient with stratified sampling. We prove that the variance of average mini-batch gradient is minimized if the elements are sampled from a weighted clustering in the gradient space. We measure the gradient variance on common deep learning benchmarks and observe that, contrary to common assumptions, gradient variance increases during training, and smaller learning rates coincide with higher variance. In addition, we introduce normalized gradient variance as a statistic that better correlates with the speed of convergence compared to gradient variance.
Forward citations
Cited by 3 Pith papers
-
VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing
VERITAS is a multi-agent system for verifiable hypothesis testing on multimodal clinical MRI datasets that achieves 81.4% verdict accuracy with frontier models and introduces an epistemic evidence labeling framework.
-
Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression
Progressive^2 improves knowledge distillation under large teacher-student capacity gaps by progressively including teacher layers and gradually compressing the student through self-distillation rounds.
-
Insights from Gradient Dynamics: Gradient Autoscaled Normalization
A hyperparameter-free gradient autoscaling method that zero-centers gradients and multiplies them by a global factor based on gradient standard deviation improves CIFAR-100 accuracy slightly on ResNets, but with weak ...
Discussion (0). Sign in to comment.