Pith. sign in

REVIEW 4 cited by

De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.09956 v3 pith:NUR4AD6D submitted 2020-02-23 cs.LG stat.ML

De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors

classification cs.LG stat.ML
keywords predictorsboundsnon-smoothdeterministicpredictordeeppac-bayesrelu-nets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In spite of several notable efforts, explaining the generalization of deterministic non-smooth deep nets, e.g., ReLU-nets, has remained challenging. Existing approaches for deterministic non-smooth deep nets typically need to bound the Lipschitz constant of such deep nets but such bounds are quite large, may even increase with the training set size yielding vacuous generalization bounds. In this paper, we present a new family of de-randomized PAC-Bayes margin bounds for deterministic non-convex and non-smooth predictors, e.g., ReLU-nets. Unlike PAC-Bayes, which applies to Bayesian predictors, the de-randomized bounds apply to deterministic predictors like ReLU-nets. A specific instantiation of the bound depends on a trade-off between the (weighted) distance of the trained weights from the initialization and the effective curvature (`flatness') of the trained predictor. To get to these bounds, we first develop a de-randomization argument for non-convex but smooth predictors, e.g., linear deep networks (LDNs), which connects the performance of the deterministic predictor with a Bayesian predictor. We then consider non-smooth predictors which for any given input realized as a smooth predictor, e.g., ReLU-nets become some LDNs for any given input, but the realized smooth predictors can be different for different inputs. For such non-smooth predictors, we introduce a new PAC-Bayes analysis which takes advantage of the smoothness of the realized predictors, e.g., LDN, for a given input, and avoids dependency on the Lipschitz constant of the non-smooth predictor. After careful de-randomization, we get a bound for the deterministic non-smooth predictor. We also establish non-uniform sample complexity results based on such bounds. Finally, we present extensive empirical results of our bounds over changing training set size and randomness in labels.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Symmetrization of Loss Functions for Robust Training of Neural Networks in the Presence of Noisy Labels

    cs.LG 2026-05 unverdicted novelty 7.0

    Symmetrization of multi-class losses produces a unique convex symmetric loss that locally approximates others and supports robust neural training under label noise.

  2. Smoothness-Based Derandomization of PAC-Bayes Bounds

    cs.LG 2026-06 unverdicted novelty 6.0

    Derives smoothness-based PAC-Bayes bounds for deterministic predictors by bounding the Jensen gap class via Rademacher complexity, yielding flatness terms in Jacobians/Hessians, and proposes a corresponding regularize...

  3. Smoothness-Based Derandomization of PAC-Bayes Bounds

    cs.LG 2026-06 unverdicted novelty 6.0

    Derives smoothness-based PAC-Bayes derandomization bounds for deterministic predictors using Rademacher complexity of the Jensen gap class, yielding Jacobian/Hessian flatness terms and a practical regularizer tested o...

  4. Symmetrization of Loss Functions for Robust Training of Neural Networks in the Presence of Noisy Labels

    cs.LG 2026-05 unverdicted novelty 6.0

    Symmetrizing cross-entropy produces the unique convex multi-class unhinged loss, which locally approximates other symmetric losses, and enables new interpolating losses SGCE and alpha-MAE with competitive performance ...