Pith. sign in

REVIEW 28 cited by

Fantastic Generalization Measures and Where to Find Them

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02178 v1 pith:R3JP2PWN submitted 2019-12-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords measuresgeneralizationnetworkscomplexitydeepexperimentsanalyzebeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiments would remain valid in other settings. We present the first large scale study of generalization in deep networks. We investigate more then 40 complexity measures taken from both theoretical bounds and empirical studies. We train over 10,000 convolutional networks by systematically varying commonly used hyperparameters. Hoping to uncover potentially causal relationships between each measure and generalization, we analyze carefully controlled experiments and show surprising failures of some measures as well as promising measures for further research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

    cs.LG 2022-01 unverdicted novelty 8.0 of 10

    Neural networks exhibit grokking on small algorithmic datasets, achieving perfect generalization well after overfitting.

  2. Pointwise Generalization in Deep Neural Networks

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Proposes pointwise Riemannian Dimension from feature eigenvalues to derive tighter, representation-aware generalization bounds for deep networks in the nonlinear regime.

  3. Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    QuBD extends algorithmic complexity estimation to quantized DNN weights, revealing that complexity decreases during learning, increases with overfitting, follows grokking patterns, and correlates with generalization.

  4. Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    HPO enables unbiased policy optimization in hybrid action spaces by mixing differentiable simulation gradients with score-function estimates, outperforming PPO as continuous dimensions increase.

  5. Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Excess risk decomposes into independent alignment (trace of inverse average Hessian times gradient covariance) and curvature terms, so both flatness and gradient alignment are required; SAGE achieves this and sets new...

  6. One task to rule them all: A closer look at traffic classification generalizability

    cs.NI 2025-07 conditional novelty 7.0 of 10

    Traffic classifiers that seem near-perfect on their own datasets fall to 30-40% accuracy on another network's same-task data, and a 1-Nearest Neighbor baseline is competitive.

  7. On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

    cs.LG 2026-08 reject novelty 6.0 of 10

    SAM's largest Hessian eigenvalue is bounded by the cube root of bGamma/(2*rho*eta^2), so larger radius, smaller batch, or larger learning rate restrict linearly stable minima to flatter regions.

  8. Characterizing Optimizer-Dependent Training Dynamics Through Hessian Eigenvector Displacement and Localization

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Hessian eigenvector displacement and inverse participation ratio metrics show SGD stabilizing leading curvature directions while Adam causes more reorganization and parameter localization in MLP training.

  9. Quantifying and Optimizing Simplicity via Polynomial Representations

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Polynomial representations yield an effective-degree simplicity metric that predicts generalization across tasks and serves as a differentiable regularizer improving performance in classification and RL.

  10. Trajectory-Based Difficulty Scoring for Reliable Learning on Tabular Data

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    TDS uses per-tree prediction trajectories to derive instance difficulty scores that rank errors better than prior hardness measures and improve active learning, selective prediction, and Mondrian conformal prediction ...

  11. Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Introduces a fairness layer for deep learning models that guarantees output parity and an online primal-dual algorithm for aggregate fairness guarantees in streaming predictions with small batch sizes.

  12. TopoGeoScore: A Self-Supervised Source-Only Geometric Framework for OOD Checkpoint Selection

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    TopoGeoScore combines a torsion-inspired Laplacian log-determinant, Ollivier-Ricci curvature, and higher-order topological summaries from source embeddings, with weights learned via self-supervised invariance to geome...

  13. Generalization at the Edge of Stability

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Training at the edge of stability causes neural network optimizers to converge on fractal attractors whose effective dimension, measured via a new sharpness dimension from the Hessian spectrum, bounds generalization e...

  14. Robust Policy Optimization to Prevent Catastrophic Forgetting

    cs.LG 2026-02 unverdicted novelty 6.0 of 10

    FRPO applies a max-min robust optimization over KL-bounded policy neighborhoods during RLHF to reduce catastrophic forgetting of safety and accuracy under subsequent SFT or RL fine-tuning.

  15. Decentralized SGD with Controlled Disagreement Finds Flatter Minima

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Keeping consensus errors alive in decentralized SGD via a learning-rate-scaled mixing term improves test accuracy and flatter minima over both DSGD and synchronous SGD.

  16. How Far Are We from True Unlearnability?

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Current unlearnable examples fail under multi-task training, and the proposed SAL and UD metrics quantify how far each method is from true unlearnability.

  17. DHEvo: Data-Algorithm Based Heuristic Evolution for Generalizable MILP Solving

    cs.NE 2025-07 conditional novelty 6.0 of 10

    DHEvo co-evolves MILP training instances and diving heuristics, improving generalization over existing LLM-based heuristic generation methods.

  18. Sharpness-Aware Minimization for Efficiently Improving Generalization

    cs.LG 2020-10 conditional novelty 6.0 of 10

    SAM solves a min-max problem to locate flat low-loss regions, improving generalization on CIFAR, ImageNet and label-noise tasks.

  19. Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization

    cs.LG 2026-07 conditional novelty 5.0 of 10

    GEAR-SAM re-allocates SAM's fixed perturbation radius across network blocks in proportion to an EMA of squared block-gradient norms, improving generalization on CIFAR, transfer, and label-noise benchmarks.

  20. Flatness Preserves Instruction Following in Vision-Language-Action Models

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    Sharpness-aware minimization during VLA finetuning preserves instruction following and yields over 60% gains across simulation and real-world tasks.

  21. SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    SingGuard presents a policy-adaptive multimodal LLM guardrail family with hybrid reasoning regimes and a new benchmark of 56,340 examples, claiming SOTA F1 across 35 datasets and improved policy adherence under runtim...

  22. SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    SingGuard introduces a policy-adaptive multimodal LLM guardrail with dynamic reasoning regimes and SingGuard-Bench, reporting SOTA F1 scores across 35 datasets and improved policy-following accuracy under runtime shifts.

  23. Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning

    cs.LG 2026-05 conditional novelty 5.0 of 10

    SAGE, an optimizer combining spectral polar-factor perturbation with gradient-agreement-scaled noise, reports 78.9% average on DomainBed.

  24. MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    MER-DG applies modality-entropy regularization to reduce fusion overfitting in multimodal domain generalization, reporting average gains of 5% over standard fusion and 2% over prior methods on EPIC-Kitchens and HAC be...

  25. An Information-Theoretic Analysis of OOD Generalization in Meta-Reinforcement Learning

    cs.LG 2025-10 unverdicted novelty 5.0 of 10

    The work establishes OOD generalization bounds for meta-supervised learning and meta-RL that exploit MDP structure, then analyzes a gradient-based meta-RL algorithm.

  26. TopoGeoScore: A Self-Supervised Source-Only Geometric Framework for OOD Checkpoint Selection

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    TopoGeoScore learns a non-negative linear combination of geometric and topological features from source embeddings via self-supervised invariance to select robust checkpoints for OOD scenarios.

  27. VASSO: Variance Suppression for Sharpness-Aware Minimization

    cs.LG 2025-09 conditional novelty 4.0 of 10

    VASSO replaces SAM's minibatch gradient with an exponential moving average of past gradients when computing the adversarial perturbation, improving generalization across vision and language tasks.

  28. Automatic Stability and Recovery for Neural Network Training

    cs.LG 2026-01 reject novelty 3.0 of 10

    A validation-loss-triggered rollback controller whose "safety guarantees" restate its own accept/reject rule.

Pith tools