Pith. sign in

hub

Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data

23 Pith papers cite this work. Polarity classification is still indexing.

23 Pith papers citing it
abstract

One of the defining properties of deep learning is that models are chosen to have many more parameters than available training data. In light of this capacity for overfitting, it is remarkable that simple algorithms like SGD reliably return solutions with low test error. One roadblock to explaining these phenomena in terms of implicit regularization, structural properties of the solution, and/or easiness of the data is that many learning bounds are quantitatively vacuous when applied to networks learned by SGD in this "deep learning" regime. Logically, in order to explain generalization, we need nonvacuous bounds. We return to an idea by Langford and Caruana (2001), who used PAC-Bayes bounds to compute nonvacuous numerical bounds on generalization error for stochastic two-layer two-hidden-unit neural networks via a sensitivity analysis. By optimizing the PAC-Bayes bound directly, we are able to extend their approach and obtain nonvacuous generalization bounds for deep stochastic neural network classifiers with millions of parameters trained on only tens of thousands of examples. We connect our findings to recent and old work on flat minima and MDL-based explanations of generalization.

hub tools

citation-role summary

background 2 method 1

citation-polarity summary

representative citing papers

Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning

cs.CV · 2026-06-08 · unverdicted · novelty 7.0

FisherAdapTune uses temporal drift in Fisher geometry, measured by scale-invariant Jensen-Shannon distance, to progressively freeze stabilized parameter groups during fine-tuning, reporting gains on segmentation and zero-shot transfer.

Pointwise Generalization in Deep Neural Networks

cs.LG · 2026-05-18 · unverdicted · novelty 7.0

Proposes pointwise Riemannian Dimension from feature eigenvalues to derive tighter, representation-aware generalization bounds for deep networks in the nonlinear regime.

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs

cs.CV · 2026-07-07 · conditional · novelty 5.0

SAMPLe adds dual gradient constraints (ERM alignment plus full-batch orthogonality) to SAM-style prompt learning and raises harmonic-mean base-to-new accuracy across CoOp, CoCoOp, MaPLe, TCP and CoPrompt.

Geometric and Information Compression of Representations in Deep Learning

cs.LG · 2026-06-19 · unverdicted · novelty 5.0

Experiments show low mutual information does not reliably correspond to geometric compression via class-wise clustering in CEB and dropout networks; the negative nonlinear link can reverse with training changes, suggesting generalization confounds the connection.

A Rigorous, Tractable Measure of Model Complexity

stat.ML · 2026-05-20 · unverdicted · novelty 5.0

A gradient-similarity complexity measure that generalizes polynomial degree, kernel length scale, neighbor count, tree splits, and forest size while offering insights into double descent.

Are Flat Minima an Illusion?

cs.LG · 2026-03-24 · conditional · novelty 5.0

The paper argues flat minima are an illusion and that a reparameterization-invariant 'weakness' score predicts generalization where raw sharpness fails.

citing papers explorer

Showing 23 of 23 citing papers.