Pith. sign in

REVIEW 2 cited by

Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04284 v5 pith:6JZ2B2CW submitted 2019-10-09 cs.LG stat.ML

classification cs.LGstat.ML
keywords marginall-layerdeepgeneralizationrobustclearoutputrelationship
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization. Unfortunately, for deep models, this relationship is less clear: existing analyses of the output margin give complicated bounds which sometimes depend exponentially on depth. In this work, we propose to instead analyze a new notion of margin, which we call the "all-layer margin." Our analysis reveals that the all-layer margin has a clear and direct relationship with generalization for deep models. This enables the following concrete applications of the all-layer margin: 1) by analyzing the all-layer margin, we obtain tighter generalization bounds for neural nets which depend on Jacobian and hidden layer norms and remove the exponential dependency on depth 2) our neural net results easily translate to the adversarially robust setting, giving the first direct analysis of robust test error for deep networks, and 3) we present a theoretically inspired training algorithm for increasing the all-layer margin. Our algorithm improves both clean and adversarially robust test performance over strong baselines in practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flat Minima and Generalization: Insights from Stochastic Convex Optimization

    cs.LG 2025-11 conditional novelty 7.0 of 10

    In smooth stochastic convex optimization, flat empirical minima can incur constant population risk while sharp minima generalize optimally, and sharpness-aware algorithms can converge to such bad flat minima.

  2. On the Sample Complexity of One Hidden Layer Networks with Equivariance, Locality and Weight Sharing

    cs.LG 2024-11 accept novelty 5.0 of 10

    For one-hidden-layer equivariant networks, generalization bounds depend only on filter norms and the sample size, while suitable weight sharing can match equivariance and locality adds an extra gain.

Pith tools