Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks

· 2015 · stat.ML · arXiv 1512.07666

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

Effective training of deep neural networks suffers from two main issues. The first is that the parameter spaces of these models exhibit pathological curvature. Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD). These methods improve convergence by adapting to the local geometry of parameter space. A second issue is overfitting, which is typically addressed by early stopping. However, recent work has demonstrated that Bayesian model averaging mitigates this problem. The posterior can be sampled by using Stochastic Gradient Langevin Dynamics (SGLD). However, the rapidly changing curvature renders default SGLD methods inefficient. Here, we propose combining adaptive preconditioners with SGLD. In support of this idea, we give theoretical properties on asymptotic convergence and predictive risk. We also provide empirical results for Logistic Regression, Feedforward Neural Nets, and Convolutional Neural Nets, demonstrating that our preconditioned SGLD method gives state-of-the-art performance on these models.

representative citing papers

Piecewise Deterministic Markov Processes for Bayesian Neural Networks

stat.ML · 2023-02-17 · unverdicted · novelty 6.0

Introduces an adaptive thinning scheme to make PDMP-based MCMC feasible for Bayesian inference in neural networks by handling model-specific IPPs efficiently.

citing papers explorer

Showing 1 of 1 citing paper.

Piecewise Deterministic Markov Processes for Bayesian Neural Networks stat.ML · 2023-02-17 · unverdicted · none · ref 7 · internal anchor
Introduces an adaptive thinning scheme to make PDMP-based MCMC feasible for Bayesian inference in neural networks by handling model-specific IPPs efficiently.

Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks

fields

years

verdicts

representative citing papers

citing papers explorer