Pith. sign in

REVIEW 1 cited by

How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11110 v1 pith:MMD7GWNM submitted 2024-06-17 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords regularizationsupportimpliciteffectfirstfunctionlayerlearn
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We investigate the ability of deep neural networks to identify the support of the target function. Our findings reveal that mini-batch SGD effectively learns the support in the first layer of the network by shrinking to zero the weights associated with irrelevant components of input. In contrast, we demonstrate that while vanilla GD also approximates the target function, it requires an explicit regularization term to learn the support in the first layer. We prove that this property of mini-batch SGD is due to a second-order implicit regularization effect which is proportional to $\eta / b$ (step size / batch size). Our results are not only another proof that implicit regularization has a significant impact on training optimization dynamics but they also shed light on the structure of the features that are learned by the network. Additionally, they suggest that smaller batches enhance feature interpretability and reduce dependency on initialization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: A Theory of Deep Learning Must Include Compositional Sparsity

    cs.LG 2025-07 conditional novelty 4.0 of 10

    All polynomial-time computable functions are compositionally sparse, and this property is the proposed reason deep networks avoid the curse of dimensionality and achieve practical success.

Pith tools