Pith. sign in

Paper Citation Record · LEDGER

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2606.06772.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06772 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T12:16:45.002429Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46861ce3-8316-4f9d-b552-4f45e5b6af44 · outbound

This paper cites A convergence theory for deep learning via over-parameterization.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent A convergence theory for deep learning via over-parameterization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.183237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.183237Z digest=sha256:6286807da545879ded5cde9efbe35ad1cfe8532683bd11820e3e6a06b15d4d1f

Observation eb492a96-147e-4c71-8017-9b8ca25ead31 · outbound

This paper cites Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.256831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.256831Z digest=sha256:d6b1109c2009c89402d3b8170d12b8b48bc7143a64833ce47fc6819f4e92f92a

Observation 5db38bcb-c9cb-4a37-b855-fae07f2aadbd · outbound

This paper cites Spectrally-normalized margin bounds for neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Spectrally-normalized margin bounds for neural networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.337823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.337823Z digest=sha256:24c7d03ea3b87279a05b7e9561f91ab714ad8e2e4b430592db0ba158a4a3e03e

Observation b544511b-b5a9-4d9d-8b9f-9f7f2512e694 · outbound

This paper cites Convergence rates for shallow neural networks learned by gradient descent.Bernoulli, 30(1):475–502, 2024.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Convergence rates for shallow neural networks learned by gradient descent.Bernoulli, 30(1):475–502, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.401888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.401888Z digest=sha256:180047282c6d0543ed207bc4d4304121a30021d54c6d2a2e3950e971fc8bdd7d

Observation e5a41aae-215a-4c0f-b9ed-ba1bd8ba6ba3 · outbound

This paper cites Stochastic Gradient Descent for Two-layer Neural Networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Stochastic Gradient Descent for Two-layer Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.496271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.496271Z digest=sha256:f7eb85c3b6e8f43e4f77ecdfe62c0122187b93db062235259b53b23b1c933420

Observation 7c7a10a3-70f3-4dfd-8b82-ac61af6ea00f · outbound

This paper cites Generalization bounds of stochastic gradient descent for wide and deep neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Generalization bounds of stochastic gradient descent for wide and deep neural networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.594800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.594800Z digest=sha256:c15116ac29a1a7ef11da9fbfb9e997271446180a6886beadf0d300240845aad6

Observation 2c5b97eb-c0b0-48cf-836a-e600d65fcca9 · outbound

This paper cites Generalization error bounds of gradient descent for learning over-parameterized deep relu networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Generalization error bounds of gradient descent for learning over-parameterized deep relu networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.659846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.659846Z digest=sha256:a5083da6e3db1f9f250354b68887a66ced8d970fdfb26f746081dc760d5fd9ad

Observation acfe2783-1236-4cbe-95cf-52467d603cae · outbound

This paper cites Optimal rates for the regularized least-squares algorithm.Foundations of Computational Mathematics, 7(3):331–368, 2007.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Optimal rates for the regularized least-squares algorithm.Foundations of Computational Mathematics, 7(3):331–368, 2007

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.725923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.725923Z digest=sha256:6471857094e067b1fb6cecab1e39866d802009ac4d9551a888434dd25ffd7032

Observation b6acc6e6-6cfc-4f73-aa2c-3134e357bb6b · outbound

This paper cites Learning with sgd and random features.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Learning with sgd and random features

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.803506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.803506Z digest=sha256:bec6aa5878e163173108b2d60f9d75fcf11b455957036b1b027ec81a10fb5d59

Observation 0e5c01a0-d39a-497f-a4bc-38d5b292c991 · outbound

This paper cites How much over-parameterization is sufficient to learn deep relu networks? InInternational Conference on Learning Representation, 2021.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent How much over-parameterization is sufficient to learn deep relu networks? InInternational Conference on Learning Representation, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.876677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.876677Z digest=sha256:b750f5f54e87d82d3edd0b175c0728193b71223cbd209ba514c1e5d8554d5b9d

Observation c4b2b81a-428b-4976-a505-9f21852e4536 · outbound

This paper cites Cambridge University Press, 2007.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Cambridge University Press, 2007

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:40.967173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:40.967173Z digest=sha256:19b85b69b481e9c34839e609d3cff9363e4c1b4bc24042e359d8cb0759cc2c62

Observation 2fdffdd1-82fc-4e57-bbb9-e5f9d0de72ad · outbound

This paper cites Nonparametric stochastic approximation with large step-sizes.Annals of Statistics, 44(4):1363–1399, 2016.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Nonparametric stochastic approximation with large step-sizes.Annals of Statistics, 44(4):1363–1399, 2016

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.032602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.032602Z digest=sha256:3667d25fdb0d95b3e5066432e507f48a8d90aa19515659404aa5a734a9c785f5

Observation 9beb56a9-124e-4ed4-bbec-9b7f6f844e6a · outbound

This paper cites Analysis of the expected $L_2$ error of an over-parametrized deep neural network estimate learned by gradient descent without regularization.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Analysis of the expected $L_2$ error of an over-parametrized deep neural network estimate learned by gradient descent without regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.098670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.098670Z digest=sha256:008572cfcb054c849be47d3f0aecf04ca1b123de3310e607cc2242885ab68ab1

Observation 585e5793-66aa-403e-a410-09ccc14e6a93 · outbound

This paper cites an unresolved cited work.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.186258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.186258Z digest=sha256:b97021a0137e65f8022310db532c548a7555919dc6953c2e93b0654eee0985ed

Observation 1e11952b-ee55-4542-ae81-d19a6e11d58a · outbound

This paper cites Gradient descent finds global minima of deep neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Gradient descent finds global minima of deep neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.302270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.302270Z digest=sha256:b7779701a13a739232631828738218ba38690f183ab90321b9722ae6326371c0

Observation 8cfc328c-d2f9-4897-8122-37587b969b79 · outbound

This paper cites Random feature amplification: Feature learning and generalization in neural networks.Journal of Machine Learning Research, 24(303):1–49, 2023.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Random feature amplification: Feature learning and generalization in neural networks.Journal of Machine Learning Research, 24(303):1–49, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.411355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.411355Z digest=sha256:59a0ab2dcb61d5393f10e35ba4a269c63802ca1f1ffa126f850684d7a8916078

Observation e876f47d-4a20-4ed0-b319-dee33b4c403f · outbound

This paper cites Size-independent sample complexity of neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Size-independent sample complexity of neural networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.419101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.419101Z digest=sha256:be6b358257a0ac226350fcdde9cd592f1af4d03b9e9a7e58beecff63bd0c976a

Observation a3853bf1-14b6-4dfd-af6a-91cb40ed6286 · outbound

This paper cites Springer Science & Business Media, 2006.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Springer Science & Business Media, 2006

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.439143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.439143Z digest=sha256:a398d8539a2d4108ebfce74dc7b8bc1a473b80cedf4ef71908f5d9e1d8fe9c0c

Observation ce3e4bc1-283d-459d-84a5-dfdc97394662 · outbound

This paper cites Deep residual learning for image recognition.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Deep residual learning for image recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.547182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.547182Z digest=sha256:4aa649a0adc17a3397bcaf5e0a2bcbedf4a7e42d6c9cc2fdb5f5281488a90c08

Observation b1d88c65-2198-4b4f-abff-9ce0c099470b · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.Advances in Neural Information Processing Systems, 31, 2018.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Neural tangent kernel: Convergence and generalization in neural networks.Advances in Neural Information Processing Systems, 31, 2018

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.635021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.635021Z digest=sha256:6d9a2d9dbd42e998f09cc786fc9e23b5408870a934c5fd949a502bd89bac6719

Observation c5c6689f-cdf6-4872-9921-8b3f1f8153cc · outbound

This paper cites Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.710849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.710849Z digest=sha256:cf765b602a660cf8b19931d777118d18826bf322848fd7e6ccd88c8e550b131e

Observation fef885de-6f29-475f-bb0d-e73c3655137d · outbound

This paper cites On the rate of convergence of an over-parametrized deep neural network regression estimate learned by gradient descent.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent On the rate of convergence of an over-parametrized deep neural network regression estimate learned by gradient descent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.764520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.764520Z digest=sha256:2fda62a50709c28d5679c3d018895bbda510f9dc90844972627ad79637c380f9

Observation e40ac79f-6146-4b2a-aa52-997e17d81182 · outbound

This paper cites Learning Lipschitz Functions by GD-trained Shallow Overparameterized ReLU Neural Networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Learning Lipschitz Functions by GD-trained Shallow Overparameterized ReLU Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.849878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.849878Z digest=sha256:426c33e6705ce096e07ed2ff314ab9ebbc58b44970eb7d631da669b0ec6a2457

Observation 1d027d6a-8464-48ae-965d-72c0df04c605 · outbound

This paper cites Stability and generalization analysis of gradient methods for shallow neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Stability and generalization analysis of gradient methods for shallow neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:41.925928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:41.925928Z digest=sha256:f7ed303d6581493bce53804c50a0e304a6ed5f2b87c1072988daa693de150605

Observation c8172789-747d-45cd-bd83-353959d5257d · outbound

This paper cites Optimization and generalization of gradient descent for shallow ReLU networks with minimal width.Journal of Machine Learning Research, 27(34):1–35, 2026.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Optimization and generalization of gradient descent for shallow ReLU networks with minimal width.Journal of Machine Learning Research, 27(34):1–35, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.003793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.003793Z digest=sha256:69ff69c159a6d5ea4c055439de113fe010a776ba6451412865969011c7fe75fc

Observation b427a89c-a4a6-4371-85a9-961882bebc09 · outbound

This paper cites Optimal rates for generalization of gradient descent for deep relu classification.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Optimal rates for generalization of gradient descent for deep relu classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.070062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.070062Z digest=sha256:a4f94782e9bc2954fedfe733e1d38de63a526bbdf1866df65849aa169b0c0068

Observation b399d351-91e6-46b4-b855-ca1fe235b8a1 · outbound

This paper cites Learning overparameterized neural networks via stochastic gradient descent on structured data.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Learning overparameterized neural networks via stochastic gradient descent on structured data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.169768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.169768Z digest=sha256:a72ec775b8d750075b3da89b0b7ef5aa693e2234fadc9fa29935f711b13c33cf

Observation 5baf53f0-844f-44fc-b0b9-f9560bcfd95d · outbound

This paper cites Optimal rates for multi-pass stochastic gradient methods.Journal of Machine Learning Research, 18(1):3375–3421, 2017.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Optimal rates for multi-pass stochastic gradient methods.Journal of Machine Learning Research, 18(1):3375–3421, 2017

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.269086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.269086Z digest=sha256:ae329702444f6963835beaca4c57a2782e10b8a026f8c6d534b992cfb0076038

Observation 299e5629-0821-4760-ab74-cd0dde57185b · outbound

This paper cites On the linearity of large non-linear models: when and why the tangent kernel is constant.Advances in Neural Information Processing Systems, 33:15954–15964, 2020.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent On the linearity of large non-linear models: when and why the tangent kernel is constant.Advances in Neural Information Processing Systems, 33:15954–15964, 2020

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.343829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.343829Z digest=sha256:56343159fbb332c2bbc4ab7e70ddfeda336a147eb008631e8d61e971a6000fc7

Observation ca49886d-529d-45c9-9a06-bc4e860c7152 · outbound

This paper cites Norm-based capacity control in neural networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Norm-based capacity control in neural networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.574038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.574038Z digest=sha256:ae82f6810bf689f2ef7d0430066184b3d3ddcab13a28e3a38910863bf0b2e7d5

Observation 046d72e2-9783-4af1-8bf2-caf656bd0e5a · outbound

This paper cites Random feature approximation for general spectral methods.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Random feature approximation for general spectral methods

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.648255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.648255Z digest=sha256:b376e9c6af33bd9343e791f588d5ee8ad68871b7e83d1c6db7ae63078dd4746d

Observation 319deed3-6fe3-428f-aff1-13b1a9720e30 · outbound

This paper cites How many neurons do we need? a refined analysis for shallow networks trained with gradient descent.Journal of Statistical Planning and Inference, page 106169, 2024.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent How many neurons do we need? a refined analysis for shallow networks trained with gradient descent.Journal of Statistical Planning and Inference, page 106169, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.842162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.842162Z digest=sha256:48e51090589b69de8544248103344caffc4bf15fb4e06f2b65da0f17144500d4

Observation f14d8b6b-c8eb-44e2-8ba6-cfc0d304c6f9 · outbound

This paper cites Gradient Descent can Learn Less Over-parameterized Two-layer Neural Networks on Classification Problems.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Gradient Descent can Learn Less Over-parameterized Two-layer Neural Networks on Classification Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:42.973908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:42.973908Z digest=sha256:67cc43c0786514550c2b29440b0562b2ddc907bb3eb6eb0ad9af2f2a2073914d

Observation f0ee101a-477d-4399-b8f3-d1ebe77c9563 · outbound

This paper cites Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.048281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.048281Z digest=sha256:c7740f459fe92a1903f1e6f8da9cec77d6d29351e5485ad71570377e5ccefb81

Observation 82cdcec2-7303-4000-9448-1b9afdb7365d · outbound

This paper cites Near-minimax optimal estimation with shallow relu neural networks.IEEE Transactions on Information Theory, 69(2):1125–1140, 2022.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Near-minimax optimal estimation with shallow relu neural networks.IEEE Transactions on Information Theory, 69(2):1125–1140, 2022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.104358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.104358Z digest=sha256:c1622931d6bf9705a829d6f85f0c23cd1dc3125611de4512c6fc37e4ab2e64f9

Observation d53033b3-6327-409c-a876-424e4ef2b89b · outbound

This paper cites Weighted sums of random kitchen sinks: Replacing minimization with random- ization in learning.Advances in Neural Information Processing Systems, 21, 2008.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Weighted sums of random kitchen sinks: Replacing minimization with random- ization in learning.Advances in Neural Information Processing Systems, 21, 2008

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.219795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.219795Z digest=sha256:8ea1835b57c691c2ada3c36974c8aec4f6cd9b9c72a98766a7d647f9aec2dfb7

Observation 9af75af2-72d0-4376-a3aa-25c8e60e56b7 · outbound

This paper cites Searching for Activation Functions.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Searching for Activation Functions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.332498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.332498Z digest=sha256:3dc40c268fa74e57ad2304950185893a5ec89d4b44a12b0b5c3f80a9a06a6dc1

Observation 61750015-719f-4098-aaa0-541dbacf7375 · outbound

This paper cites Stability & generalisation of gradient descent for shallow neural networks without the neural tangent kernel.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Stability & generalisation of gradient descent for shallow neural networks without the neural tangent kernel

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.424537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.424537Z digest=sha256:a14170b65c9824d85804388a7630afd1db349d4d3fafbaab263b98cbfc69d4e8

Observation a1c0171e-1882-441a-84f7-b7c8994e6b72 · outbound

This paper cites Springer Science & Business Media, 2008.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Springer Science & Business Media, 2008

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.576741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.576741Z digest=sha256:5a9656544467f939983b7368f9def4c4e03a8c77f184cfa186e074453fe44166

Observation 0f661f08-e358-4d5b-aa11-f656006170bd · outbound

This paper cites Generalization and stability of interpolating neural networks with minimal width.Journal of Machine Learning Research, 25(156):1–41, 2024.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Generalization and stability of interpolating neural networks with minimal width.Journal of Machine Learning Research, 25(156):1–41, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.680524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.680524Z digest=sha256:8066398e3d1eb3064210b552efd075f7f30bb905e1e2557fc892f0c4a05df24b

Observation 2a749880-26d4-4cfe-81ff-a68e8cdbc3b4 · outbound

This paper cites Sharper guarantees for learning neural network classifiers with gradient methods, 2025.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Sharper guarantees for learning neural network classifiers with gradient methods, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.824107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.824107Z digest=sha256:c18e656f5d43c9f41796b8fff88bb3377306c324527a8061e2519d17e6a8bc70

Observation 076a1d6f-2328-4b00-901d-d52b835415c1 · outbound

This paper cites Cambridge university press, 2018.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Cambridge university press, 2018

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:43.945052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:43.945052Z digest=sha256:57da3fb525e0dcd96676a1baffb65b7af4a500985df5c260b7c557cd9c5c4d87

Observation 19b07b83-10f3-407c-a37e-2b46e3462c4a · outbound

This paper cites Cambridge university press, 2019.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Cambridge university press, 2019

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.084254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.084254Z digest=sha256:f2d0a0616044c79e51e60676c2abd85b8b6f6f28d8e30c188dfc8ff8e628a1ce

Observation 7239a9fa-7c11-426d-81af-3bb2967b83ae · outbound

This paper cites Generalization guarantees of gradient descent for shallow neural networks.Neural Computation, 37(2):344–402, 2025.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Generalization guarantees of gradient descent for shallow neural networks.Neural Computation, 37(2):344–402, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.187228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.187228Z digest=sha256:d2946eca6830c84203481f833dadf29413bdb1150414cbbc88502c836f3376f0

Observation 9e3af01a-01a2-472a-a250-691a01565b73 · outbound

This paper cites Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.264407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.264407Z digest=sha256:d0487a6e3f1320106db690a0e7d933fe1555657babacc4ebddd12810ed375737

Observation c1567103-02cb-460b-a37d-a13222c32d1a · outbound

This paper cites Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov-Arnold Networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov-Arnold Networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.426478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.426478Z digest=sha256:3a068cf01e6184c4ca7631c4a8781b759192835e4a4e4e19352ee70c3d1074b8

Observation 706e4076-2479-4362-a672-143a68964aec · outbound

This paper cites an unresolved cited work.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.584752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.584752Z digest=sha256:24b7bb390db50ebe047be792a3d7ef13b9d414629ac12cfae0006b0f8b92085c

Observation de22d84f-9daf-4271-bf36-b230ddf6756f · outbound

This paper cites Learning bounds for kernel regression using effective data dimensionality.Neural computation, 17(9):2077–2098, 2005.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Learning bounds for kernel regression using effective data dimensionality.Neural computation, 17(9):2077–2098, 2005

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.693702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.693702Z digest=sha256:c04fa7452e994ca6b42bff9937762cd49cc282f5260a950bd4ec1c945e635bf0

Observation c8d4b679-ad0a-4bfb-9242-6cf57697f373 · outbound

This paper cites Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:44.835136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:44.835136Z digest=sha256:e74a2080f9b0c5af77ff3de502b984bfc9ac2ad15524c612ff787961dae6382a

Observation e96c0f94-8b89-4aea-b143-92746457fe19 · outbound

This paper cites Their equation (47) is guaranteed by Lemma 18 with κ2 =∥K∥ ∞, Γ =n , δ=δ 2, ζi =K xi, Q= R X Kx ⊗K xdρx.

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent Their equation (47) is guaranteed by Lemma 18 with κ2 =∥K∥ ∞, Γ =n , δ=δ 2, ζi =K xi, Q= R X Kx ⊗K xdρx

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T12:16:45.002429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:16:45.002429Z digest=sha256:246d8f6748859a64c544180d91839cafc16169b296d776cc2633574e1672ce0d

Pith citing papers

No inbound Pith citation observations are available.