Pith. sign in

hub

Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits

14 Pith papers cite this work, alongside 18 external citations. Polarity classification is still indexing.

14 Pith papers citing it
18 external citations · Pith
abstract

Weight averaging of Stochastic Gradient Descent (SGD) iterates is a popular method for training deep learning models. While it is often used as part of complex training pipelines to improve generalization or serve as a `teacher' model, weight averaging lacks proper evaluation on its own. In this work, we present a systematic study of the Exponential Moving Average (EMA) of weights. We first explore the training dynamics of EMA, give guidelines for hyperparameter tuning, and highlight its good early performance, partly explaining its success as a teacher. We also observe that EMA requires less learning rate decay compared to SGD since averaging naturally reduces noise, introducing a form of implicit regularization. Through extensive experiments, we show that EMA solutions differ from last-iterate solutions. EMA models not only generalize better but also exhibit improved i) robustness to noisy labels, ii) prediction consistency, iii) calibration and iv) transfer learning. Therefore, we suggest that an EMA of weights is a simple yet effective plug-in to improve the performance of deep learning models.

hub tools

citation-role summary

background 3 method 1

citation-polarity summary

years

2026 13 2025 1

representative citing papers

Stabilizing distribution-free probabilistic forecasts

cs.LG · 2026-05-27 · unverdicted · novelty 7.0

Neural network-parameterized regression splines enable joint optimization of forecast quality and stability in distribution-free probabilistic time series models by penalizing dissimilarities from forecast updates.

Exploring Line Bundle Standard Models with Transformers

hep-th · 2026-06-30 · conditional · novelty 5.0

A Transformer trained by reinforcement learning generates heterotic line-bundle sums that satisfy anomaly-cancellation, stability, and chirality constraints, and its policy transfers usefully across Calabi-Yau geometries.

Refined Analysis of Entropy-Regularized Actor-Critic

cs.LG · 2026-05-23 · unverdicted · novelty 5.0

Exact critic in entropy-regularized actor-critic yields strong variance reduction, enabling Õ(log(1/ε)) sample complexity for ε-optimal regularized value.

citing papers explorer

Showing 14 of 14 citing papers.