Sensitivity and Generalization in Neural Networks: an Empirical Study

Daniel A. Abolafia; Jascha Sohl-Dickstein; Jeffrey Pennington; Roman Novak; Yasaman Bahri

arxiv: 1802.08760 · v3 · pith:ZOQKHEORnew · submitted 2018-02-23 · 📊 stat.ML · cs.AI· cs.LG· cs.NE

Sensitivity and Generalization in Neural Networks: an Empirical Study

Roman Novak , Yasaman Bahri , Daniel A. Abolafia , Jeffrey Pennington , Jascha Sohl-Dickstein This is my paper

classification 📊 stat.ML cs.AIcs.LGcs.NE

keywords generalizationcomplexitynetworksneuralassociateddataempiricalfactors

0 comments

read the original abstract

In practice it is often found that large over-parameterized neural networks generalize better than their smaller counterparts, an observation that appears to conflict with classical notions of function complexity, which typically favor smaller models. In this work, we investigate this tension between complexity and generalization through an extensive empirical exploration of two natural metrics of complexity related to sensitivity to input perturbations. Our experiments survey thousands of models with various fully-connected architectures, optimizers, and other hyper-parameters, as well as four different image classification datasets. We find that trained neural networks are more robust to input perturbations in the vicinity of the training data manifold, as measured by the norm of the input-output Jacobian of the network, and that it correlates well with generalization. We further establish that factors associated with poor generalization $-$ such as full-batch training or using random labels $-$ correspond to lower robustness, while factors associated with good generalization $-$ such as data augmentation and ReLU non-linearities $-$ give rise to more robust functions. Finally, we demonstrate how the input-output Jacobian norm can be predictive of generalization at the level of individual test points.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Geometric Asymmetry in MoE Specialization: Functional Decorrelation and Representational Overlap
cs.LG 2026-05 unverdicted novelty 7.0

MoE experts in pretrained Transformers exhibit functional decorrelation with near-zero Jacobian alignment yet occupy partially overlapping representation subspaces, with routing sparsity modulating the geometry.
Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems
cs.LG 2026-05 unverdicted novelty 6.0

A new scale-aware diagnostic framework shows that unconstrained diffusion generative models exhibit structural freezing and instability instead of smooth physical responses under multiscale perturbations.
Complexity of Linear Regions in Self-supervised Deep ReLU Networks
cs.LG 2026-04 unverdicted novelty 6.0

Self-supervised ReLU networks form substantially fewer linear regions than supervised models for comparable accuracy, with contrastive methods rapidly expanding regions and self-distillation consolidating them, enabli...
Escape dynamics and implicit bias of one-pass SGD in overparameterized quadratic networks
cond-mat.dis-nn 2026-04 unverdicted novelty 6.0

In overparameterized quadratic networks, one-pass SGD escapes generalization plateaus only modestly faster and selects the initialization-closest zero-loss solution due to a conserved quantity in the overlap ODEs.