Pith. sign in

REVIEW 4 cited by

Kernel and Rich Regimes in Overparametrized Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.05827 v3 pith:FUF42SSW submitted 2019-06-13 cs.LG stat.ML

Kernel and Rich Regimes in Overparametrized Models

classification cs.LG stat.ML
keywords kernelrichmodelsmultilayernetworksoverparametrizedregimestransition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effect of finding the minimum RKHS norm solution. This stands in contrast to other studies which demonstrate how gradient descent on overparametrized multilayer networks can induce rich implicit biases that are not RKHS norms. Building on an observation by Chizat and Bach, we show how the scale of the initialization controls the transition between the "kernel" (aka lazy) and "rich" (aka active) regimes and affects generalization properties in multilayer homogeneous models. We provide a complete and detailed analysis for a simple two-layer model that already exhibits an interesting and meaningful transition between the kernel and rich regimes, and we demonstrate the transition for more complex matrix factorization models and multilayer non-linear networks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Estimation of High Dimensional Bounded Discrete Graphical Models via Regularized Generalized Score Matching

    stat.ME 2026-06 unverdicted novelty 6.0

    Introduces bounded discrete graphical models and the BRIDGE regularized score matching estimator with nonasymptotic error bounds and exact support recovery for high-dimensional discrete data.

  2. A Theory on Flow Matching with Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0

    Establishes convergence guarantees for overparameterized 2-layer ReLU networks in flow matching, generalization bounds for the velocity-field objective, and Wasserstein guarantees for generated samples, using multi-ta...

  3. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

    cs.LG 2024-01 unverdicted novelty 6.0

    SPIN lets weak LLMs become strong by self-generating training data from previous model versions and training to prefer human-annotated responses over its own outputs, outperforming DPO even with extra GPT-4 data on be...

  4. Domain Adaptation of Mismatched Proximal Denoiser for Plug-and-Play Image Reconstruction

    eess.IV 2026-07 conditional novelty 5.0

    For PnP-PGD, residual reconstruction error is bounded by average squared mismatch between the deployed denoiser and the target proximal map, motivating proximal-matching few-shot adaptation that outperforms MSE adapta...