Pith. sign in

Liberty or Depth: Deep Bayesian Neural Nets Do Not Need Complex Weight Posterior Approximations

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We challenge the longstanding assumption that the mean-field approximation for variational inference in Bayesian neural networks is severely restrictive, and show this is not the case in deep networks. We prove several results indicating that deep mean-field variational weight posteriors can induce similar distributions in function-space to those induced by shallower networks with complex weight posteriors. We validate our theoretical contributions empirically, both through examination of the weight posterior using Hamiltonian Monte Carlo in small models and by comparing diagonal- to structured-covariance in large settings. Since complex variational posteriors are often expensive and cumbersome to implement, our results suggest that using mean-field variational inference in a deeper model is both a practical and theoretically justified alternative to structured approximations.

citation-role summary

background 1

citation-polarity summary

fields

stat.ML 1

years

2019 1

verdicts

ACCEPT 1

roles

background 1

polarities

background 1

representative citing papers

On the Expressiveness of Approximate Inference in Bayesian Neural Networks

stat.ML · 2019-09-02 · accept · novelty 8.0

For single-hidden-layer ReLU Bayesian neural networks, mean-field Gaussian and Monte Carlo dropout posteriors provably cannot express higher predictive variance between well-separated low-variance regions, and this limitation persists empirically in deep networks despite a universality theorem.

citing papers explorer

Showing 1 of 1 citing paper.

  • On the Expressiveness of Approximate Inference in Bayesian Neural Networks stat.ML · 2019-09-02 · accept · none · ref 10 · internal anchor

    For single-hidden-layer ReLU Bayesian neural networks, mean-field Gaussian and Monte Carlo dropout posteriors provably cannot express higher predictive variance between well-separated low-variance regions, and this limitation persists empirically in deep networks despite a universality theorem.