Pith. sign in

REVIEW 17 cited by

User-friendly introduction to PAC-Bayes bounds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.11216 v6 pith:MDNCOIHI submitted 2021-10-21 stat.ML cs.LGmath.STstat.TH

User-friendly introduction to PAC-Bayes bounds

classification stat.ML cs.LGmath.STstat.TH
keywords boundspac-bayespredictorsdistributionintroductionprobabilitysomeaccording
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Aggregated predictors are obtained by making a set of basic predictors vote according to some weights, that is, to some probability distribution. Randomized predictors are obtained by sampling in a set of basic predictors, according to some prescribed probability distribution. Thus, aggregated and randomized predictors have in common that they are not defined by a minimization problem, but by a probability distribution on the set of predictors. In statistical learning theory, there is a set of tools designed to understand the generalization ability of such procedures: PAC-Bayesian or PAC-Bayes bounds. Since the original PAC-Bayes bounds of D. McAllester, these tools have been considerably improved in many directions (we will for example describe a simplified version of the localization technique of O. Catoni that was missed by the community, and later rediscovered as "mutual information bounds"). Very recently, PAC-Bayes bounds received a considerable attention: for example there was workshop on PAC-Bayes at NIPS 2017, "(Almost) 50 Shades of Bayesian Learning: PAC-Bayesian trends and insights", organized by B. Guedj, F. Bach and P. Germain. One of the reason of this recent success is the successful application of these bounds to neural networks by G. Dziugaite and D. Roy. An elementary introduction to PAC-Bayes theory is still missing. This is an attempt to provide such an introduction.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Statistical Cost of Adaptation in Multi-Source Transfer Learning

    math.ST 2026-05 unverdicted novelty 8.0

    Multi-source transfer learning incurs an intrinsic adaptation cost that can exceed one, with phase transitions separating regimes where bias-agnostic estimators match oracle performance from those where they cannot.

  2. Post-Cut Metadata Inference Attacks on Quantum Circuit Cutting Pipelines

    quant-ph 2026-04 conditional novelty 8.0

    Post-cut metadata from quantum circuit fragments enables high-accuracy inference of algorithm family, cut mechanism, and Hamiltonian structure via machine learning on fragment width, depth, and gate counts.

  3. Aggregation with Exponential Weights is Optimal in Expectation

    math.ST 2026-07 unverdicted novelty 7.0

    AEW achieves the excess risk bound T log(M)/(n+1) in expectation for sufficiently large constant temperature T under i.i.d. random design and bounded L-Lipschitz μ-strongly convex losses.

  4. Sample Complexity of Scientific Discovery: PAC Learnability of Compositional Function Trees

    cs.LG 2026-06 unverdicted novelty 7.0

    Proves that Rademacher complexity of depth-d compositional trees over finite operator vocabulary is controlled by (K b L)^{d} / sqrt(n) under Lipschitz conditions on operators.

  5. Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data

    cs.LG 2026-05 unverdicted novelty 7.0

    ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.

  6. Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data

    cs.LG 2026-05 unverdicted novelty 7.0

    Asymmetric Langevin Unlearning uses public data to suppress unlearning noise costs by O(1/n_pub²), enabling practical mass unlearning with preserved utility under distribution mismatch.

  7. Post-Cut Metadata Inference Attacks on Quantum Circuit Cutting Pipelines

    quant-ph 2026-04 unverdicted novelty 7.0

    Circuit-cutting metadata leaks algorithm family and Hamiltonian k-locality with near-perfect accuracy via topological transpilation penalties on production QPUs.

  8. Tighter Information-Theoretic Generalization Bounds via a Novel Class of Change of Measure Inequalities

    cs.IT 2026-02 conditional novelty 7.0

    A unified data-processing framework produces tighter change-of-measure inequalities that improve information-theoretic generalization bounds across learning theory and privacy.

  9. Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation Regime

    cs.LG 2025-10 conditional novelty 7.0

    New PAC-Bayes bounds for the Gibbs posterior remain non-vacuous in the interpolation regime and can be approximated by Langevin Monte Carlo, but the tight experimental numbers rely on an unproved random-label calibrat...

  10. PAC-Bayesian Certificates for Quadratic Closed-Loop Control

    eess.SY 2026-06 unverdicted novelty 6.0

    PAC-Bayesian bounds are derived for quadratic closed-loop control via SLS parameterization, yielding Chernoff certificates for posteriors over responses, a mean-response deployment result, and a data-driven learning a...

  11. On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

    cs.LG 2026-06 unverdicted novelty 6.0

    Prompt-conditioned LLMs face irreducible error floors from language's limited information capacity and alignment constraints, proven via PAC-Bayes bounds on bilevel cheap-talk games for certain task families.

  12. PAC-Bayes Bounds for Gibbs Posteriors via Singular Learning Theory

    stat.ML 2026-04 unverdicted novelty 6.0

    PAC-Bayes bounds for Gibbs posteriors are obtained via singular learning theory, producing explicit and tighter posterior-averaged risk bounds that adapt to data structure in overparameterized models.

  13. Distributionally Robust PAC-Bayesian Control

    cs.LG 2026-04 unverdicted novelty 6.0

    A distributionally robust PAC-Bayesian approach derives sub-Gaussian loss proxies and performance bounds tied to closed-loop operator norms via system level synthesis, enabling optimization-based safety certificates f...

  14. Federated Learning with Nonvacuous Generalisation Bounds

    cs.LG 2023-10 unverdicted novelty 6.0

    Federated learning trains private local randomised predictors whose aggregation yields a global predictor with nonvacuous PAC-Bayesian generalisation bounds and near-centralized accuracy.

  15. Change of measure through the Legendre transform

    stat.ML 2022-02 unverdicted novelty 6.0

    Derives f-divergence change-of-measure inequalities via Legendre transform to extend PAC-Bayes bounds beyond the classical Donsker-Varadhan setting.

  16. Towards the Connection between Activation Sparsity and Flat Minima

    cs.LG 2026-05 unverdicted novelty 5.0

    MLP activation sparsity equals augmented flatness divided by input norm times gradient; the ratio falls during training and can be reduced further by three plug-and-play changes, yielding higher sparsity on ImageNet and C4.

  17. Tighter Information-Theoretic Generalization Bounds via a Novel Class of Change of Measure Inequalities

    cs.IT 2026-02 conditional novelty 5.0

    A unified DPI-based framework yields novel change-of-measure inequalities that produce tighter high-probability generalization bounds.