Pith. sign in

REVIEW 2 cited by

Understanding Why Neural Networks Generalize Well Through GSNR of Parameters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.07384 v2 pith:UAZSIIDK submitted 2020-01-21 cs.LG stat.ML

classification cs.LGstat.ML
keywords gsnrdnnsduringgeneralizationgradientparameterstraininggeneralize
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As deep neural networks (DNNs) achieve tremendous success across many application domains, researchers tried to explore in many aspects on why they generalize well. In this paper, we provide a novel perspective on these issues using the gradient signal to noise ratio (GSNR) of parameters during training process of DNNs. The GSNR of a parameter is defined as the ratio between its gradient's squared mean and variance, over the data distribution. Based on several approximations, we establish a quantitative relationship between model parameters' GSNR and the generalization gap. This relationship indicates that larger GSNR during training process leads to better generalization performance. Moreover, we show that, different from that of shallow models (e.g. logistic regression, support vector machines), the gradient descent optimization dynamics of DNNs naturally produces large GSNR during training, which is probably the key to DNNs' remarkable generalization ability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LUViT jointly pretrains a ViT with masked auto-encoding and LoRA adapters in a frozen LLM block, reporting +0.4% ImageNet-1K accuracy and up to +2.2% on ImageNet-A over its own MAE baseline.

  2. DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DeepKD is a knowledge distillation trainer that decouples task, target-class, and non-target-class gradients with GSNR-based momentum and a dynamic top-k mask, yielding consistent accuracy gains on CIFAR-100, ImageNet...

Pith tools