Pith. sign in

REVIEW 2 cited by

Are All Layers Created Equal?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.01996 v4 pith:6ULPEM5D submitted 2019-02-06 stat.ML cs.AIcs.LG

Are All Layers Created Equal?

classification stat.ML cs.AIcs.LG
keywords layersdeepmodelsnetworksrobustcriticalevidenceexperimental
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Understanding deep neural networks is a major research objective with notable experimental and theoretical attention in recent years. The practical success of excessively large networks underscores the need for better theoretical analyses and justifications. In this paper we focus on layer-wise functional structure and behavior in overparameterized deep models. To do so, we study empirically the layers' robustness to post-training re-initialization and re-randomization of the parameters. We provide experimental results which give evidence for the heterogeneity of layers. Morally, layers of large deep neural networks can be categorized as either "robust" or "critical". Resetting the robust layers to their initial values does not result in adverse decline in performance. In many cases, robust layers hardly change throughout training. In contrast, re-initializing critical layers vastly degrades the performance of the network with test error essentially dropping to random guesses. Our study provides further evidence that mere parameter counting or norm calculations are too coarse in studying generalization of deep models, and "flatness" and robustness analysis of trained models need to be examined while taking into account the respective network architectures.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization

    cs.LG 2026-07 conditional novelty 6.0

    An information-theoretic 'local redundancy' is lower-bounded, via an entropy-cancellation argument, by the expected squared gradient norm on synthetic probe data, and this proxy modestly out-predicts existing plastici...

  2. Domain Adaptation of Mismatched Proximal Denoiser for Plug-and-Play Image Reconstruction

    eess.IV 2026-07 conditional novelty 5.0

    For PnP-PGD, residual reconstruction error is bounded by average squared mismatch between the deployed denoiser and the target proximal map, motivating proximal-matching few-shot adaptation that outperforms MSE adapta...