Pith. sign in

REVIEW 3 cited by

Stop Wasting My Time! Saving Days of ImageNet and BERT Training with Latest Weight Averaging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.14981 v2 pith:LS6BLTX7 submitted 2022-09-29 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords trainingaveragingdaysimagenetlatestmodeltimeweights
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training vision or language models on large datasets can take days, if not weeks. We show that averaging the weights of the k latest checkpoints, each collected at the end of an epoch, can speed up the training progression in terms of loss and accuracy by dozens of epochs, corresponding to time savings up to ~68 and ~30 GPU hours when training a ResNet50 on ImageNet and RoBERTa-Base model on WikiText-103, respectively. We also provide the code and model checkpoint trajectory to reproduce the results and facilitate research on reusing historical weights for faster convergence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 4 citations worldwide. Full citation record

  1. EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes

    cs.LG 2025-07 unverdicted novelty 5.0 of 10

    A bias-corrected exponential moving average (BEMA) is claimed to remove the lag of standard EMA weight averaging during LLM fine-tuning, improving convergence and final performance over EMA and vanilla training.

  2. WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Checkpoint merging during constant-LR training can replace LR decay and yields improved LLM benchmark scores over Warmup-Stable-Decay.

  3. SeWA: Selective Weight Average via Probabilistic Masking

    cs.LG 2025-02 conditional novelty 5.0 of 10

    SeWA adaptively selects a sparse set of checkpoints for weight averaging via learned probabilistic masks, claiming better generalization with fewer averaged points.

Pith tools