Proposes pointwise Riemannian Dimension from feature eigenvalues to derive tighter, representation-aware generalization bounds for deep networks in the nonlinear regime.
Deep Learning is Not So Mysterious or Different
9 Pith papers cite this work. Polarity classification is still indexing.
abstract
Deep neural networks are often seen as different from other model classes by defying conventional notions of generalization. Popular examples of anomalous generalization behaviour include benign overfitting, double descent, and the success of overparametrization. We argue that these phenomena are not distinct to neural networks, or particularly mysterious. Moreover, this generalization behaviour can be intuitively understood, and rigorously characterized, using long-standing generalization frameworks such as PAC-Bayes and countable hypothesis bounds. We present soft inductive biases as a key unifying principle in explaining these phenomena: rather than restricting the hypothesis space to avoid overfitting, embrace a flexible hypothesis space, with a soft preference for simpler solutions that are consistent with the data. This principle can be encoded in many model classes, and thus deep learning is not as mysterious or different from other model classes as it might seem. However, we also highlight how deep learning is relatively distinct in other ways, such as its ability for representation learning, phenomena such as mode connectivity, and its relative universality.
citation-role summary
citation-polarity summary
representative citing papers
Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
Spectral Born machines are Fourier-phase quantum generative models over Z_d^n that train classically via graph-spectral MMD and show reduced parameters plus apparent overfitting resistance on integer data.
Predictively consistent priors let complex Bayesian models match or beat the out-of-sample performance of selected simpler models across linear, logistic, and nonlinear examples without explicit selection.
A LightGBM classifier trained on NWAY Bayesian matches identifies true Chandra-Gaia counterparts for 113k X-ray sources, flags 7k ambiguous cases, and attributes half of 20k separation-only matches to chance coincidences, validated at 95% on COUP without positional features.
A model-independent upper bound on generalization gap is established that depends solely on the Rényi entropy of the data-generating distribution for histogram-determined algorithms such as ERM.
Expanding feature capacity enables discovery of sparse priced-risk structures, so nonlinear expansions plus basis pursuit outperform ridgeless methods beyond a complexity threshold.
Quantum computers may enable more natural manipulation of Fourier spectra in ML models via the Quantum Fourier Transform, potentially leading to resource-efficient spectral methods.
Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.
citing papers explorer
-
Pointwise Generalization in Deep Neural Networks
Proposes pointwise Riemannian Dimension from feature eigenvalues to derive tighter, representation-aware generalization bounds for deep networks in the nonlinear regime.
-
Estimating Implicit Regularization in Deep Learning
Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
-
Spectral Born machines: classically trainable quantum generative models for discrete data
Spectral Born machines are Fourier-phase quantum generative models over Z_d^n that train classically via graph-spectral MMD and show reduced parameters plus apparent overfitting resistance on integer data.
-
To select or not to select: predictively consistent priors instead of model selection
Predictively consistent priors let complex Bayesian models match or beat the out-of-sample performance of selected simpler models across linear, logistic, and nonlinear examples without explicit selection.
-
The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning
A LightGBM classifier trained on NWAY Bayesian matches identifies true Chandra-Gaia counterparts for 113k X-ray sources, flags 7k ambiguous cases, and attributes half of 20k separation-only matches to chance coincidences, validated at 95% on COUP without positional features.
-
Overfitting has a limitation: a model-independent generalization gap bound based on R\'enyi entropy
A model-independent upper bound on generalization gap is established that depends solely on the Rényi entropy of the data-generating distribution for histogram-determined algorithms such as ERM.
-
The Virtue of Sparsity in Complexity
Expanding feature capacity enables discovery of sparse priced-risk structures, so nonlinear expansions plus basis pursuit outperform ridgeless methods beyond a complexity threshold.
-
Spectral methods: crucial for machine learning, natural for quantum computers?
Quantum computers may enable more natural manipulation of Fourier spectra in ML models via the Quantum Fourier Transform, potentially leading to resource-efficient spectral methods.
-
Statistical Properties of Training & Generalization
Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.