Pith. sign in

REVIEW 5 cited by

Deep Equals Shallow for ReLU Networks in Kernel Regimes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.14397 v4 pith:7TYGLYOG submitted 2020-09-30 stat.ML cs.LG

classification stat.MLcs.LG
keywords deepkernelnetworksshallowapproximationkernelscertaineigenvalue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep networks are often considered to be more expressive than shallow ones in terms of approximation. Indeed, certain functions can be approximated by deep networks provably more efficiently than by shallow ones, however, no tractable algorithms are known for learning such deep models. Separately, a recent line of work has shown that deep networks trained with gradient descent may behave like (tractable) kernel methods in a certain over-parameterized regime, where the kernel is determined by the architecture and initialization, and this paper focuses on approximation for such kernels. We show that for ReLU activations, the kernels derived from deep fully-connected networks have essentially the same approximation properties as their shallow two-layer counterpart, namely the same eigenvalue decay for the corresponding integral operator. This highlights the limitations of the kernel framework for understanding the benefits of such deep architectures. Our main theoretical result relies on characterizing such eigenvalue decays through differentiability properties of the kernel function, which also easily applies to the study of other kernels defined on the sphere.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. Querying Kernel Methods Suffices for Reconstructing their Training Data

    cs.LG 2025-05 conditional novelty 8.0 of 10

    Query-only access to kernel regression, SVM and KDE models suffices to reconstruct their exact training points, via a measure-theoretic proof and image experiments.

  2. The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation

    math.ST 2026-07 accept novelty 7.0 of 10

    Matching minimax prediction rates for discretely observed functional linear regression are n^{-ν/(ν+1)}+(nm)^{-ν/κ} under independent design, and those two terms plus m^{-ν}+m^{-4α} under common design.

  3. Variable Importance Identification Through Lazy Training for Binary Classification

    stat.ML 2026-07 reject novelty 6.0 of 10

    LazyVI tests feature importance in binary classifiers by fitting a linearized logistic model on neural tangent features and claims op(n^{-1/2}) asymptotic normality.

  4. Revisiting the Neural Tangent Kernel: the role of large width and depth

    cs.LG 2025-11 reject novelty 5.0 of 10

    For infinite-width ReLU nets, the normalized neural tangent kernel provably collapses toward the all-ones matrix with depth, while the claimed convergence of the kernel-regression solution to a nontrivial limit is onl...

  5. On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations

    cs.LG 2025-08 reject novelty 4.0 of 10

    The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.

Pith tools