Pith. sign in

REVIEW 6 cited by

Every Model Learned by Gradient Descent Is Approximately a Kernel Machine

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.00152 v1 pith:4LYCZG2F submitted 2020-11-30 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords kernellearningdeepapproximatelydatadescentfunctiongradient
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep learning's successes are often attributed to its ability to automatically discover new representations of the data, rather than relying on handcrafted features like other learning methods. We show, however, that deep networks learned by the standard gradient descent algorithm are in fact mathematically approximately equivalent to kernel machines, a learning method that simply memorizes the data and uses it directly for prediction via a similarity function (the kernel). This greatly enhances the interpretability of deep network weights, by elucidating that they are effectively a superposition of the training examples. The network architecture incorporates knowledge of the target function into the kernel. This improved understanding should lead to better learning algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 49 citations worldwide. Full citation record

  1. Adaptive kernel predictors from feature-learning infinite limits of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).

  2. Region-wise stacking ensembles for estimating brain-age using MRI

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A region-wise stacking ensemble improves brain age prediction from structural MRI compared to regional mean gray matter volume.

  3. A Comprehensive Framework for Automated Segmentation of Perivascular Spaces in Brain MRI with the nnU-Net

    eess.IV 2024-11 conditional novelty 6.0 of 10

    A voxel-spacing agnostic nnU-Net, trained with sparse annotations, iterative label cleaning, and 12,740 pseudo-labelled images, reaches DSC 85.6% for perivascular space segmentation and is extended to midbrain, hippoc...

  4. Multiplex Nodal Modularity: A novel network metric for the regional analysis of amnestic mild cognitive impairment during a working memory binding task

    q-bio.NC 2025-01 reject novelty 4.0 of 10

    A per-node decomposition of multiplex modularity is applied to working-memory binding fMRI and DTI networks and is claimed to flag MCI-to-AD converters.

  5. A Brain Age Residual Biomarker (BARB): Leveraging MRI-Based Models to Detect Latent Health Conditions in U.S. Veterans

    cs.LG 2025-01 reject novelty 4.0 of 10

    A CNN ensemble predicts age from veteran MRIs with R2 0.816, but the evidence that residuals flag chronic disease is weakened by post hoc analysis and age confounding.

  6. Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL

    cs.LG 2025-06 conditional novelty 3.0 of 10

    Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.

Pith tools