Pith. sign in

REVIEW 7 cited by

Generalization in Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.05468 v9 pith:5QERG2OR submitted 2017-10-16 stat.ML cs.AIcs.LGcs.NE

Generalization in Deep Learning

classification stat.ML cs.AIcs.LGcs.NE
keywords deeplearningdiscussgeneralizationopentheoreticalalgorithmicapproaches
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper provides theoretical insights into why and how deep learning can generalize well, despite its large capacity, complexity, possible algorithmic instability, nonrobustness, and sharp minima, responding to an open question in the literature. We also discuss approaches to provide non-vacuous generalization guarantees for deep learning. Based on theoretical observations, we propose new open problems and discuss the limitations of our results.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Dangerous Liaisons of Convex Learning and Non-Affine Aggregation

    cs.LG 2026-06 unverdicted novelty 8.0

    Monotonicity of aggregated gradients holds if and only if the aggregation rule is positively affine; non-affine rules therefore prevent steady convergence and degrade stability.

  2. Deep Learning Scaling is Predictable, Empirically

    cs.LG 2017-12 unverdicted novelty 7.0

    Deep learning generalization error follows power-law scaling with training set size across multiple domains, with model size scaling sublinearly with data size.

  3. Adaptive Kernel Ridge Regression with Linear Structure: Sharp Oracle Inequalities and Minimax Optimality

    math.ST 2026-05 unverdicted novelty 6.0

    An augmented kernel ridge regression estimator separates linear and nonlinear components to achieve sharp oracle inequalities and minimax optimal prediction risk under general kernels.

  4. WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval

    cs.CV 2026-04 unverdicted novelty 5.0

    WRF4CIR uses weight-regularized fine-tuning with adversarial perturbations to mitigate overfitting in composed image retrieval and narrows the generalization gap on benchmarks.

  5. Generalization error bounds for two-layer neural networks with Lipschitz loss function

    stat.ML 2026-04 unverdicted novelty 4.0

    Generalization error bounds of order O(n^{-1/2}) (dimension-free) are derived for two-layer neural networks with Lipschitz losses under independent test data, and O(n^{-1/(d_in + d_out)}) without independence, using W...

  6. DNNs, Dataset Statistics, and Correlation Functions

    physics.hist-ph 2025-11 unverdicted novelty 4.0

    DNNs succeed by capturing high-order correlation structures in datasets, similar to mesoscale methods in physics.

  7. Sharpness-Aware Minimization with Z-Score Gradient Filtering

    cs.LG 2025-05 unverdicted novelty 4.0

    Z-Score Filtered SAM retains only high absolute Z-score gradient components per layer during the ascent step and reports higher test accuracy than standard SAM on CIFAR and Tiny-ImageNet benchmarks.