Pith. sign in

REVIEW 3 cited by

Local Kernel Renormalization as a mechanism for feature learning in overparametrized Convolutional Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.11807 v1 pith:TLZE3MQZ submitted 2023-07-21 cs.LG cond-mat.dis-nn

classification cs.LGcond-mat.dis-nn
keywords networksfeaturelearningkernelarchitecturesconvolutionalfinite-widthneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Feature learning, or the ability of deep neural networks to automatically learn relevant features from raw data, underlies their exceptional capability to solve complex tasks. However, feature learning seems to be realized in different ways in fully-connected (FC) or convolutional architectures (CNNs). Empirical evidence shows that FC neural networks in the infinite-width limit eventually outperform their finite-width counterparts. Since the kernel that describes infinite-width networks does not evolve during training, whatever form of feature learning occurs in deep FC architectures is not very helpful in improving generalization. On the other hand, state-of-the-art architectures with convolutional layers achieve optimal performances in the finite-width regime, suggesting that an effective form of feature learning emerges in this case. In this work, we present a simple theoretical framework that provides a rationale for these differences, in one hidden layer networks. First, we show that the generalization performance of a finite-width FC network can be obtained by an infinite-width network, with a suitable choice of the Gaussian priors. Second, we derive a finite-width effective action for an architecture with one convolutional hidden layer and compare it with the result available for FC networks. Remarkably, we identify a completely different form of kernel renormalization: whereas the kernel of the FC architecture is just globally renormalized by a single scalar parameter, the CNN kernel undergoes a local renormalization, meaning that the network can select the local components that will contribute to the final prediction in a data-dependent way. This finding highlights a simple mechanism for feature learning that can take place in overparametrized shallow CNNs, but not in shallow FC architectures or in locally connected neural networks without weight sharing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation

    stat.ML 2025-10 conditional novelty 8.0 of 10

    A replica/HCIZ theory predicts the Bayes-optimal generalization error of proportional-width MLPs near interpolation and discovers layer-wise specialization transitions that make deeper targets harder to learn.

  2. Proportional infinite-width infinite-depth limit for deep linear neural networks

    stat.ML 2024-11 accept novelty 7.0 of 10

    Deep linear neural networks in the proportional depth-width limit converge to a nontrivial mixture of Gaussians, with posterior output correlations that depend on the observed labels.

  3. Kernel shape renormalization explains output-output correlations in finite Bayesian one-hidden-layer networks

    cond-mat.dis-nn 2024-12 conditional novelty 4.0 of 10

    Output-output correlations in finite Bayesian one-hidden-layer networks follow the kernel shape renormalization order parameter, with readout weight overlap equal to Q*_ab/λ1.

Pith tools