Pith. sign in

REVIEW 3 cited by

Post-mortem on a deep learning contest: a Simpson's paradox and the complementary roles of scale metrics versus shape metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.00734 v2 pith:D72DBSIF submitted 2021-06-01 cs.LG stat.ML

classification cs.LGstat.ML
keywords modelsmetricstheorygeneralizationhyperparametersmetricperformwhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To understand better good generalization performance in state-of-the-art neural network (NN) models, and in particular the success of the ALPHAHAT metric based on Heavy-Tailed Self-Regularization (HT-SR) theory, we analyze of a corpus of models that was made publicly-available for a contest to predict the generalization accuracy of NNs. These models include a wide range of qualities and were trained with a range of architectures and regularization hyperparameters. We break ALPHAHAT into its two subcomponent metrics: a scale-based metric; and a shape-based metric. We identify what amounts to a Simpson's paradox: where "scale" metrics (from traditional statistical learning theory) perform well in aggregate, but can perform poorly on subpartitions of the data of a given depth, when regularization hyperparameters are varied; and where "shape" metrics (from HT-SR theory) perform well on each subpartition of the data, when hyperparameters are varied for models of a given depth, but can perform poorly overall when models with varying depths are aggregated. Our results highlight the subtlety of comparing models when both architectures and hyperparameters are varied; the complementary role of implicit scale versus implicit shape parameters in understanding NN model quality; and the need to go beyond one-size-fits-all metrics based on upper bounds from generalization theory to describe the performance of NN models. Our results also clarify further why the ALPHAHAT metric from HT-SR theory works so well at predicting generalization across a broad range of CV and NLP models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SETOL: A Semi-Empirical Theory of (Deep) Learning

    cs.LG 2025-07 conditional novelty 7.0 of 10

    SETOL derives the HTSR layer quality metrics as integrated R-transforms of the layer spectral density, and proposes a determinant condition (ERG) as a marker of ideal learning.

  2. Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

    cs.LG 2025-09 conditional novelty 6.0 of 10

    For diagonal and quadratic two-layer networks, training maps to LASSO and matrix compressed sensing, yielding a full phase diagram of excess-risk scaling exponents and a spectral characterization of the trained weights.

  3. Models of Heavy-Tailed Mechanistic Universality

    stat.ML 2025-06 conditional novelty 6.0 of 10

    A new random matrix model with one structure parameter explains heavy-tailed spectra in trained networks, and yields scaling laws, optimizer-tail behavior, and a description of the five-plus-one phases of training.

Pith tools