REVIEW 17 cited by
User-friendly introduction to PAC-Bayes bounds
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
User-friendly introduction to PAC-Bayes bounds
read the original abstract
Aggregated predictors are obtained by making a set of basic predictors vote according to some weights, that is, to some probability distribution. Randomized predictors are obtained by sampling in a set of basic predictors, according to some prescribed probability distribution. Thus, aggregated and randomized predictors have in common that they are not defined by a minimization problem, but by a probability distribution on the set of predictors. In statistical learning theory, there is a set of tools designed to understand the generalization ability of such procedures: PAC-Bayesian or PAC-Bayes bounds. Since the original PAC-Bayes bounds of D. McAllester, these tools have been considerably improved in many directions (we will for example describe a simplified version of the localization technique of O. Catoni that was missed by the community, and later rediscovered as "mutual information bounds"). Very recently, PAC-Bayes bounds received a considerable attention: for example there was workshop on PAC-Bayes at NIPS 2017, "(Almost) 50 Shades of Bayesian Learning: PAC-Bayesian trends and insights", organized by B. Guedj, F. Bach and P. Germain. One of the reason of this recent success is the successful application of these bounds to neural networks by G. Dziugaite and D. Roy. An elementary introduction to PAC-Bayes theory is still missing. This is an attempt to provide such an introduction.
Forward citations
Cited by 17 Pith papers
-
The Statistical Cost of Adaptation in Multi-Source Transfer Learning
Multi-source transfer learning incurs an intrinsic adaptation cost that can exceed one, with phase transitions separating regimes where bias-agnostic estimators match oracle performance from those where they cannot.
-
Post-Cut Metadata Inference Attacks on Quantum Circuit Cutting Pipelines
Post-cut metadata from quantum circuit fragments enables high-accuracy inference of algorithm family, cut mechanism, and Hamiltonian structure via machine learning on fragment width, depth, and gate counts.
-
Aggregation with Exponential Weights is Optimal in Expectation
AEW achieves the excess risk bound T log(M)/(n+1) in expectation for sufficiently large constant temperature T under i.i.d. random design and bounded L-Lipschitz μ-strongly convex losses.
-
Sample Complexity of Scientific Discovery: PAC Learnability of Compositional Function Trees
Proves that Rademacher complexity of depth-d compositional trees over finite operator vocabulary is controlled by (K b L)^{d} / sqrt(n) under Lipschitz conditions on operators.
-
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.
-
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
Asymmetric Langevin Unlearning uses public data to suppress unlearning noise costs by O(1/n_pub²), enabling practical mass unlearning with preserved utility under distribution mismatch.
-
Post-Cut Metadata Inference Attacks on Quantum Circuit Cutting Pipelines
Circuit-cutting metadata leaks algorithm family and Hamiltonian k-locality with near-perfect accuracy via topological transpilation penalties on production QPUs.
-
Tighter Information-Theoretic Generalization Bounds via a Novel Class of Change of Measure Inequalities
A unified data-processing framework produces tighter change-of-measure inequalities that improve information-theoretic generalization bounds across learning theory and privacy.
-
Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation Regime
New PAC-Bayes bounds for the Gibbs posterior remain non-vacuous in the interpolation regime and can be approximated by Langevin Monte Carlo, but the tight experimental numbers rely on an unproved random-label calibrat...
-
PAC-Bayesian Certificates for Quadratic Closed-Loop Control
PAC-Bayesian bounds are derived for quadratic closed-loop control via SLS parameterization, yielding Chernoff certificates for posteriors over responses, a mean-response deployment result, and a data-driven learning a...
-
On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners
Prompt-conditioned LLMs face irreducible error floors from language's limited information capacity and alignment constraints, proven via PAC-Bayes bounds on bilevel cheap-talk games for certain task families.
-
PAC-Bayes Bounds for Gibbs Posteriors via Singular Learning Theory
PAC-Bayes bounds for Gibbs posteriors are obtained via singular learning theory, producing explicit and tighter posterior-averaged risk bounds that adapt to data structure in overparameterized models.
-
Distributionally Robust PAC-Bayesian Control
A distributionally robust PAC-Bayesian approach derives sub-Gaussian loss proxies and performance bounds tied to closed-loop operator norms via system level synthesis, enabling optimization-based safety certificates f...
-
Federated Learning with Nonvacuous Generalisation Bounds
Federated learning trains private local randomised predictors whose aggregation yields a global predictor with nonvacuous PAC-Bayesian generalisation bounds and near-centralized accuracy.
-
Change of measure through the Legendre transform
Derives f-divergence change-of-measure inequalities via Legendre transform to extend PAC-Bayes bounds beyond the classical Donsker-Varadhan setting.
-
Towards the Connection between Activation Sparsity and Flat Minima
MLP activation sparsity equals augmented flatness divided by input norm times gradient; the ratio falls during training and can be reduced further by three plug-and-play changes, yielding higher sparsity on ImageNet and C4.
-
Tighter Information-Theoretic Generalization Bounds via a Novel Class of Change of Measure Inequalities
A unified DPI-based framework yields novel change-of-measure inequalities that produce tighter high-probability generalization bounds.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.