Pith. sign in

REVIEW 4 cited by

Deep Ensembles Secretly Perform Empirical Bayes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17917 v1 pith:AIZLYK6L submitted 2025-01-29 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords ensemblesdeepbayesianempiricalpriorbnnslearnedbayes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Quantifying uncertainty in neural networks is a highly relevant problem which is essential to many applications. The two predominant paradigms to tackle this task are Bayesian neural networks (BNNs) and deep ensembles. Despite some similarities between these two approaches, they are typically surmised to lack a formal connection and are thus understood as fundamentally different. BNNs are often touted as more principled due to their reliance on the Bayesian paradigm, whereas ensembles are perceived as more ad-hoc; yet, deep ensembles tend to empirically outperform BNNs, with no satisfying explanation as to why this is the case. In this work we bridge this gap by showing that deep ensembles perform exact Bayesian averaging with a posterior obtained with an implicitly learned data-dependent prior. In other words deep ensembles are Bayesian, or more specifically, they implement an empirical Bayes procedure wherein the prior is learned from the data. This perspective offers two main benefits: (i) it theoretically justifies deep ensembles and thus provides an explanation for their strong empirical performance; and (ii) inspection of the learned prior reveals it is given by a mixture of point masses -- the use of such a strong prior helps elucidate observed phenomena about ensembles. Overall, our work delivers a newfound understanding of deep ensembles which is not only of interest in it of itself, but which is also likely to generate future insights that drive empirical improvements for these models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Ray Tracing Sampler: Bayesian Sampling of Neural Networks for Everyone

    astro-ph.IM 2025-10 conditional novelty 7.0 of 10

    A new ray-tracing MCMC sampler keeps ray speed constant, making it far more robust to stochastic gradients and able to sample billion-parameter neural networks on one GPU.

  2. ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Combining LoRA with snapshot ensembling yields a parameter-efficient uncertainty-aware segmentation ensemble that matches snapshot full-rank baselines, with feed-forward layers identified as the critical LoRA target.

  3. Last Layer Empirical Bayes

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Last-layer empirical Bayes trains a normalizing-flow prior on the final layer by maximizing expected log-likelihood, and matches but does not beat existing uncertainty quantification methods.

  4. Bayesian Neural Networks versus deep ensembles for uncertainty quantification in machine learning interatomic potentials

    physics.chem-ph 2025-09 conditional novelty 4.0 of 10

    Deep ensembles outperform variational Bayesian neural networks in accuracy and uncertainty calibration on machine-learned TiO2 potentials, with some low-data exceptions when Bayesian methods are better calibrated.

Pith tools