Pith. sign in

REVIEW 2 cited by

Bayesian Neural Network Priors Revisited

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.06571 v3 pith:EKS6DQIM submitted 2021-02-12 stat.ML cs.LG

classification stat.MLcs.LG
keywords priorsnetworkneuralbayesiancolddisplaydistributionseffect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Isotropic Gaussian priors are the de facto standard for modern Bayesian neural network inference. However, it is unclear whether these priors accurately reflect our true beliefs about the weight distributions or give optimal performance. To find better priors, we study summary statistics of neural network weights in networks trained using stochastic gradient descent (SGD). We find that convolutional neural network (CNN) and ResNet weights display strong spatial correlations, while fully connected networks (FCNNs) display heavy-tailed weight distributions. We show that building these observations into priors can lead to improved performance on a variety of image classification datasets. Surprisingly, these priors mitigate the cold posterior effect in FCNNs, but slightly increase the cold posterior effect in ResNets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Student's t likelihood (ν=5) is a robust default for VI-trained BNNs, improving CRPS in most tested settings while occasionally losing on MSE to Gaussian under lognormal noise.

  2. ALAS: Additive Learnable Alpha-Stable Kernels for Flexible Bayesian Optimization

    cs.LG 2026-06 conditional novelty 4.0 of 10

    ALAS learns the spectral tail exponent of a GP kernel, adapting from Gaussian to heavy-tailed behavior, with a per-dimension additive variant for high-dimensional Bayesian optimization.

Pith tools