Pith. sign in

REVIEW 2 cited by

Neural Kernels Without Tangents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.02237 v2 pith:6GTZOQUV submitted 2020-03-04 cs.LG stat.ML

Neural Kernels Without Tangents

classification cs.LG stat.ML
keywords neuralkernelscompositionalkernelnetworksaccuracyachievesblocks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We investigate the connections between neural networks and simple building blocks in kernel space. In particular, using well established feature space tools such as direct sum, averaging, and moment lifting, we present an algebra for creating "compositional" kernels from bags of features. We show that these operations correspond to many of the building blocks of "neural tangent kernels (NTK)". Experimentally, we show that there is a correlation in test error between neural network architectures and the associated kernels. We construct a simple neural network architecture using only 3x3 convolutions, 2x2 average pooling, ReLU, and optimized with SGD and MSE loss that achieves 96% accuracy on CIFAR10, and whose corresponding compositional kernel achieves 90% accuracy. We also use our constructions to investigate the relative performance of neural networks, NTKs, and compositional kernels in the small dataset regime. In particular, we find that compositional kernels outperform NTKs and neural networks outperform both kernel methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Domain Adaptation of Mismatched Proximal Denoiser for Plug-and-Play Image Reconstruction

    eess.IV 2026-07 conditional novelty 5.0

    For PnP-PGD, residual reconstruction error is bounded by average squared mismatch between the deployed denoiser and the target proximal map, motivating proximal-matching few-shot adaptation that outperforms MSE adapta...

  2. Interpretable Self-Supervised Learning via Representer Landmarks and Nystr\"om Approximation

    cs.LG 2025-09 reject novelty 5.0

    KREPES trains kernel models for self-supervised learning at scale with Nyström landmarks and reads out interpretability directly from the learned coefficients.