Pith. sign in

REVIEW 1 cited by

A self consistent theory of Gaussian Processes captures feature learning effects in finite CNNs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.04110 v1 pith:7ZB2QCKM submitted 2021-06-08 cs.LG cond-mat.stat-mechstat.ML

classification cs.LGcond-mat.stat-mechstat.ML
keywords learningfeaturednnseffectsanalyticalconsistentdeepfinite
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural networks (DNNs) in the infinite width/channel limit have received much attention recently, as they provide a clear analytical window to deep learning via mappings to Gaussian Processes (GPs). Despite its theoretical appeal, this viewpoint lacks a crucial ingredient of deep learning in finite DNNs, laying at the heart of their success -- feature learning. Here we consider DNNs trained with noisy gradient descent on a large training set and derive a self consistent Gaussian Process theory accounting for strong finite-DNN and feature learning effects. Applying this to a toy model of a two-layer linear convolutional neural network (CNN) shows good agreement with experiments. We further identify, both analytical and numerically, a sharp transition between a feature learning regime and a lazy learning regime in this model. Strong finite-DNN effects are also derived for a non-linear two-layer fully connected network. Our self consistent theory provides a rich and versatile analytical framework for studying feature learning and other non-lazy effects in finite DNNs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Width-Robust Learnability in Mean-Field Bayesian Neural Networks

    stat.ML 2026-07 conditional novelty 7.0 of 10

    For fixed-depth mean-field Bayesian nets on the Boolean cube, poly-sample learnability at infinite width equals poly-width learnability equals poly-bounded reduced entropy.

Pith tools