Pith. sign in

REVIEW 1 cited by

When Do Flat Minima Optimizers Work?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00661 v5 pith:CFGOUXUO submitted 2022-02-01 cs.LG stat.ML

classification cs.LGstat.ML
keywords optimizersacrossbeenbenchmarkingimprovelearningstochasticadaptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, flat-minima optimizers, which seek to find parameters in low-loss neighborhoods, have been shown to improve a neural network's generalization performance over stochastic and adaptive gradient-based optimizers. Two methods have received significant attention due to their scalability: 1. Stochastic Weight Averaging (SWA), and 2. Sharpness-Aware Minimization (SAM). However, there has been limited investigation into their properties and no systematic benchmarking of them across different domains. We fill this gap here by comparing the loss surfaces of the models trained with each method and through broad benchmarking across computer vision, natural language processing, and graph representation learning tasks. We discover several surprising findings from these results, which we hope will help researchers further improve deep learning optimizers, and practitioners identify the right optimizer for their problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weight Averaging for Out-of-Distribution Generalization and Few-Shot Domain Adaptation

    cs.CV 2025-01 reject novelty 4.0 of 10

    Gradient-similarity-regularized weight averaging and WA+SAM fine-tuning are tested on OOD and few-shot domain adaptation benchmarks, with mixed results that do not support the claimed improvements.

Pith tools