Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs

Andrew Gordon Wilson; Dmitrii Podoprikhin; Dmitry Vetrov; Pavel Izmailov; Timur Garipov

arxiv: 1802.10026 · v4 · pith:V7MWSMFJnew · submitted 2018-02-27 · 📊 stat.ML · cs.AI· cs.LG

Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs

Timur Garipov , Pavel Izmailov , Dmitrii Podoprikhin , Dmitry Vetrov , Andrew Gordon Wilson This is my paper

classification 📊 stat.ML cs.AIcs.LG

keywords ensemblinggeometriclosscomplexensemblesfastfunctionstrain

0 comments

read the original abstract

The loss functions of deep neural networks are complex and their geometric properties are not well understood. We show that the optima of these complex loss functions are in fact connected by simple curves over which training and test accuracy are nearly constant. We introduce a training procedure to discover these high-accuracy pathways between modes. Inspired by this new geometric insight, we also propose a new ensembling method entitled Fast Geometric Ensembling (FGE). Using FGE we can train high-performing ensembles in the time required to train a single model. We achieve improved performance compared to the recent state-of-the-art Snapshot Ensembles, on CIFAR-10, CIFAR-100, and ImageNet.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol
cs.LG 2026-05 unverdicted novelty 7.0

Introduces a template-controlled difference-in-differences protocol that corrects chat-template confounding when measuring alignment-induced activation shifts in LLMs and recovers the refusal direction with higher fidelity.
Recoverable but Not Stationary:Local Linear Structures in Weights and Activations
cs.LG 2026-06 unverdicted novelty 6.0

Local low-rank task-gradient structures exist in weights and activations but are non-stationary, with initial recovery updates forming a basis capturing 77% of LoRA displacement and parameter steps aligning 0.58 cosin...
Otter Weather: Skillful and Computationally Efficient Medium-Range Weather Forecasting
cs.LG 2026-06 unverdicted novelty 5.0

Otter Weather is a spatiotemporal model that outperforms NWP baselines by 9.6% at 24h lead with under 3.5 A100-days training and extends efficiency gains to probabilistic forecasting via CRPS.
Statistical Properties of Training & Generalization
stat.ML 2026-06 unverdicted novelty 2.0

Neural scaling laws in deep learning interact with physics constraints and inductive biases beyond classical statistics.