Pith. sign in

REVIEW 9 cited by

Fantastic Generalization Measures and Where to Find Them

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02178 v1 pith:R3JP2PWN submitted 2019-12-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords measuresgeneralizationnetworkscomplexitydeepexperimentsanalyzebeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiments would remain valid in other settings. We present the first large scale study of generalization in deep networks. We investigate more then 40 complexity measures taken from both theoretical bounds and empirical studies. We train over 10,000 convolutional networks by systematically varying commonly used hyperparameters. Hoping to uncover potentially causal relationships between each measure and generalization, we analyze carefully controlled experiments and show surprising failures of some measures as well as promising measures for further research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    SAGE, an optimizer combining spectral polar-factor perturbation with gradient-agreement-scaled noise, reports 78.9% average on DomainBed.

  2. One task to rule them all: A closer look at traffic classification generalizability

    cs.NI 2025-07 conditional novelty 7.0 of 10

    Traffic classifiers that seem near-perfect on their own datasets fall to 30-40% accuracy on another network's same-task data, and a 1-Nearest Neighbor baseline is competitive.

  3. On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

    cs.LG 2026-08 reject novelty 6.0 of 10

    SAM's largest Hessian eigenvalue is bounded by the cube root of bGamma/(2*rho*eta^2), so larger radius, smaller batch, or larger learning rate restrict linearly stable minima to flatter regions.

  4. Decentralized SGD with Controlled Disagreement Finds Flatter Minima

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Keeping consensus errors alive in decentralized SGD via a learning-rate-scaled mixing term improves test accuracy and flatter minima over both DSGD and synchronous SGD.

  5. How Far Are We from True Unlearnability?

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Current unlearnable examples fail under multi-task training, and the proposed SAL and UD metrics quantify how far each method is from true unlearnability.

  6. DHEvo: Data-Algorithm Based Heuristic Evolution for Generalizable MILP Solving

    cs.NE 2025-07 conditional novelty 6.0 of 10

    DHEvo co-evolves MILP training instances and diving heuristics, improving generalization over existing LLM-based heuristic generation methods.

  7. Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization

    cs.LG 2026-07 conditional novelty 5.0 of 10

    GEAR-SAM re-allocates SAM's fixed perturbation radius across network blocks in proportion to an EMA of squared block-gradient norms, improving generalization on CIFAR, transfer, and label-noise benchmarks.

  8. VASSO: Variance Suppression for Sharpness-Aware Minimization

    cs.LG 2025-09 conditional novelty 4.0 of 10

    VASSO replaces SAM's minibatch gradient with an exponential moving average of past gradients when computing the adversarial perturbation, improving generalization across vision and language tasks.

  9. Automatic Stability and Recovery for Neural Network Training

    cs.LG 2026-01 reject novelty 3.0 of 10

    A validation-loss-triggered rollback controller whose "safety guarantees" restate its own accept/reject rule.

Pith tools