Exact finite-sample analysis reveals that the PPC-based selector for the learning rate η is η-invariant with known variance and flat prior, and data-independent with unknown variance and reference prior, collapsing to the smallest grid value before any data are observed.
Bayesian inference for the learning rate in Generalised Bayesian inference
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In Generalised Bayesian Inference (GBI), the learning rate and hyperparameters of the loss must be estimated. These inference-hyperparameters can't be estimated jointly with the other parameters, from the data, by giving them a prior. However, in some settings there exist unknown ``true'' hyperparameter-values about which it is meaningful to have prior belief. It is then possible to use Bayesian inference with held-out data to get hyperparameter-posteriors. We define two hyperparameter posteriors, one based on an ELPPD-utility and one aiming to cover the pseudo-true parameter. The new framework supports estimation and uncertainty quantification for multiple hyperparameters jointly. Experiments show that the resulting GBI-posteriors out-perform Bayesian inference on simulated test data and select optimal or near optimal hyperparameter values in a large real problem of text analysis. Generalised Bayesian inference is particularly useful for combining multiple data sets and most of our examples belong to that setting. We also give asymptotic results for some of the special ``multi-modular'' Generalised Bayes posteriors which we use in our examples.
fields
stat.ME 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence
Exact finite-sample analysis reveals that the PPC-based selector for the learning rate η is η-invariant with known variance and flat prior, and data-independent with unknown variance and reference prior, collapsing to the smallest grid value before any data are observed.