REVIEW 2 major objections 2 minor 72 references
Sharp Sobolev Sandwich and Approximation Rates of Radon-Domain $L^p$ Ridge Integral Spaces for ReLU$^k$ Networks
T0 review · 2 major / 2 minor · reviewed 2026-06-25 · grok-4.3
Pith's one-line read The Radon-domain L^p space of ridge integrals recovers the critical Sobolev space H^{k+(d+1)/2} exactly when p=2.
desk verdict The paper defines a Radon-domain L^p space for ReLU^k ridge integrals, proves it recovers the critical Sobolev space at p=2 via Fourier analysis, and extracts explicit approximation rates from that link. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Radon-domain L^p space R L^p_k(Ω) consisting of ridge-integral representations whose coefficient densities belong to L^p in the Radon domain; the space encodes the combined effect of activation regularity k and Radon back-projection on Sobolev regularity.
What would settle it
Exhibit a concrete function in H^{k+(d+1)/2}(Ω) whose Radon-domain coefficient density fails to belong to L^2, or compute the exact Sobolev index shift produced by the Radon transform multiplier and check whether it matches the Seeger-Sogge-Stein amount.
Extended reading notes
Core claim
The Radon-domain L^p space R L^p_k(Ω) recovers the critical Sobolev space H^{k+(d+1)/2}(Ω) for p=2 by elementary Fourier analysis, while for 1<p<∞ it forms a Sobolev sandwich whose gap on each side equals the Seeger-Sogge-Stein loss for the Radon transform.
Load-bearing premise
That every function in the target Sobolev space admits a ridge integral representation whose coefficient density lies in the required L^p space on the Radon domain.
Editorial extensions
If this is right
- The identification yields the optimal Hilbert-space approximation rate O(n^{-1/2 - (2k+1)/(2d)}) for linearized ReLU^k networks at p=2.
- Discretization of the ridge integral by a deterministic interpolation skeleton plus uniform sampling produces high-probability L^p approximation rates for any 1<p<∞.
- The joint action of the activation power k and the Radon-transform loss fixes the precise Sobolev regularity attainable by the network class.
Reading between the lines
- The same Radon-domain construction could be applied to other integral representations to obtain Sobolev characterizations for networks with different activations.
- Iterating the ridge-integral representation might give analogous sandwich results for deeper networks.
- The explicit loss term suggests that quadrature rules adapted to the Radon geometry could improve practical training rates beyond generic sampling.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Radon-domain L^p space ΡL^p_k(Ω) consisting of functions on a bounded domain Ω that admit a ridge-integral representation with L^p coefficient density in the Radon domain. It claims that for p=2 this space coincides exactly with the critical Sobolev space H^{k+(d+1)/2}(Ω) via elementary Fourier analysis, while for 1<p<∞ the space is sandwiched between two Sobolev spaces whose gap on each side equals the Seeger–Sogge–Stein loss of the Radon transform as a Fourier integral operator. The paper further discretizes the representation via deterministic interpolation and uniform sampling to obtain high-probability L^p approximation rates, including the optimal rate O(n^{-1/2-(2k+1)/(2d)}) for linearized networks at p=2.
Significance. If the identifications and rates are rigorously established, the work supplies a sharp functional-analytic characterization of the approximation spaces realized by shallow ReLU^k networks, directly tying activation smoothness to Sobolev regularity through the Radon back-projection. The explicit discretization yielding optimal Hilbert-space rates constitutes a concrete, falsifiable contribution.
major comments (2)
- [§3] §3 (Hilbert-space identification): The claimed equality ΡL^2_k(Ω) = H^{k+(d+1)/2}(Ω) is asserted via global Fourier analysis and the Fourier-slice theorem for the Radon transform, yet the manuscript supplies no explicit extension operator E: H^{k+(d+1)/2}(Ω) o H^{k+(d+1)/2}(R^d) that preserves the existence of an L^2 Radon-domain density. Without such an operator the restriction argument does not close, and the equivalence may fail or acquire boundary-dependent constants.
- [§4] §4 (Sobolev sandwich): The statement that the gap on each side of the sandwich is exactly the Seeger–Sogge–Stein loss is load-bearing for the sharpness claim, but the proof sketch does not verify that the Radon-domain L^p norm controls the precise loss term after localization to the bounded domain Ω; a concrete estimate relating the two norms is required.
minor comments (2)
- [Introduction] Notation: the symbol ΡL^p_k(Ω) is introduced without an explicit comparison table to the classical ridge-function spaces used in the neural-network literature; adding one would improve readability.
- [Discretization section] The discretization argument in the final section invokes uniform sampling but does not state the precise probability space or the measure on the Radon domain; a short paragraph clarifying the sampling measure would remove ambiguity.
Simulated Author's Rebuttal
We thank the referee for the careful reading and constructive major comments. Both points identify places where the manuscript would benefit from additional explicit constructions and estimates; we will incorporate these in the revision.
read point-by-point responses
-
Referee: [§3] §3 (Hilbert-space identification): The claimed equality ΡL^2_k(Ω) = H^{k+(d+1)/2}(Ω) is asserted via global Fourier analysis and the Fourier-slice theorem for the Radon transform, yet the manuscript supplies no explicit extension operator E: H^{k+(d+1)/2}(Ω) → H^{k+(d+1)/2}(R^d) that preserves the existence of an L^2 Radon-domain density. Without such an operator the restriction argument does not close, and the equivalence may fail or acquire boundary-dependent constants.
Authors: We agree that an explicit extension operator must be supplied to close the argument rigorously. In the revised manuscript we will insert a dedicated paragraph (or short subsection) constructing a bounded linear extension E: H^{k+(d+1)/2}(Ω) → H^{k+(d+1)/2}(R^d) that is compatible with the Fourier-slice theorem; one standard choice is the Stein extension operator (or a reflection-based extension when Ω is Lipschitz), whose Fourier multiplier properties ensure that the Radon-domain L^2 density remains in L^2 after extension. With this operator the global Fourier identification on R^d restricts correctly to Ω, yielding the claimed equality with constants independent of the particular extension. revision: yes
-
Referee: [§4] §4 (Sobolev sandwich): The statement that the gap on each side of the sandwich is exactly the Seeger–Sogge–Stein loss is load-bearing for the sharpness claim, but the proof sketch does not verify that the Radon-domain L^p norm controls the precise loss term after localization to the bounded domain Ω; a concrete estimate relating the two norms is required.
Authors: We accept that a concrete localization estimate is required. In the revision we will add an explicit lemma that, for a smooth cutoff χ supported in a neighborhood of Ω, relates the Radon-domain L^p norm of χf to the Seeger–Sogge–Stein loss term plus a controllable remainder. The proof proceeds by writing the localized Radon transform as a Fourier integral operator, applying the known global loss bounds, and estimating the commutator terms arising from the cutoff via integration by parts and the smoothness of the Radon kernel; the resulting constants depend only on the diameter of Ω and the support of χ, thereby confirming that the sandwich gap is precisely the Seeger–Sogge–Stein loss after localization. revision: yes
Circularity Check
No circularity: Fourier identification and rates are self-contained
full rationale
The paper's central identification of RL^2_k(Ω) with H^{k+(d+1)/2}(Ω) is obtained directly via elementary Fourier analysis applied to the ridge-integral representation; the Sobolev sandwich for 1<p<∞ invokes the external Seeger-Sogge-Stein loss for the Radon transform as a Fourier integral operator. Neither step reduces to a fitted parameter, self-definition, or self-citation chain. The approximation rates follow from explicit discretization of the integral representation via deterministic interpolation and uniform sampling, producing concrete high-probability bounds without circular reduction to the input data or definitions. The analysis is self-contained against external benchmarks.
Assumptions & free parameters
assumptions (2)
- standard math Elementary Fourier analysis suffices to recover the Sobolev regularity from the ridge integral representation
- domain assumption The Radon transform acts as a Fourier integral operator whose loss is given exactly by the Seeger-Sogge-Stein theorem
invented entities (1)
-
Radon-domain L^p space R L^p_k(Ω)
Cite this review
Pith. "Pith review of Sharp Sobolev Sandwich and Approximation Rates of Radon-Domain $L^p$ Ridge Integral Spaces for ReLU$^k$ Networks." pith.science (2026). https://pith.science/paper/SKK7L7BJ
@misc{pith2026260624795,
author = {Pith},
title = {Pith review of: Sharp Sobolev Sandwich and Approximation Rates of Radon-Domain $L^p$ Ridge Integral Spaces for ReLU$^k$ Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/SKK7L7BJ}},
note = {Machine review of arXiv:2606.24795}
}
abstract
We develop the $L^p$ space and approximation theory for shallow neural networks with $\mathrm{ReLU}^k$ activations. The central object is the Radon-domain $L^p$ space $\mathcal{R}L^p_k(\Omega)$ containing all functions on a bounded domain $\Omega$ that admit a ridge integral representation whose coefficient density belongs to $L^p$ in the Radon domain. In the Hilbert case $p=2$, we prove by elementary Fourier analysis that this space recovers the critical Sobolev space $H^{k+(d+1)/2}(\Omega)$. For general $1<p<\infty$, the identity becomes a Sobolev sandwich. The sharp gap of each side is exactly the Seeger--Sogge--Stein loss for the Radon transform as a Fourier integral operator. This also clarifies how the activation regularity and Radon back-projection jointly produce the regularity. As an application, we discretize the integral representation using a deterministic interpolation skeleton plus uniform sampling. This yields high-probability $L^p$ approximation rates and the optimal Hilbert rate $O\!\big(n^{-\frac12-\frac{2k+1}{2d}}\big)$ at $p=2$ for linearized neural networks.
Figures
Reference graph
Works this paper leans on
-
[1]
Adams and John J
Robert A. Adams and John J. F. Fournier.Sobolev Spaces. Number 140 in Pure and Applied Mathematics. Academic Press, Amsterdam Heidelberg, 2. ed., reprinted edition, 2008
2008
-
[2]
Kalton.Topics in Banach Space Theory, volume 233 ofGraduate Texts in Mathematics
Fernando Albiac and Nigel J. Kalton.Topics in Banach Space Theory, volume 233 ofGraduate Texts in Mathematics. Springer-Verlag, New York, 2006
2006
-
[3]
Breaking the Curse of Dimensionality with Convex Neural Net- works.Journal of Machine Learning Research, 18(19):1–53, 2017
Francis Bach. Breaking the Curse of Dimensionality with Convex Neural Net- works.Journal of Machine Learning Research, 18(19):1–53, 2017
2017
-
[4]
On the Equivalence between Kernel Quadrature Rules and Ran- dom Feature Expansions.Journal of Machine Learning Research, 18(21):1–38, 2017
Francis Bach. On the Equivalence between Kernel Quadrature Rules and Ran- dom Feature Expansions.Journal of Machine Learning Research, 18(21):1–38, 2017
2017
-
[5]
A.R. Barron. Universal Approximation Bounds for Superpositions of a Sigmoidal Function.IEEE Transactions on Information Theory, 39(3):930–945, May 1993
1993
-
[6]
Understanding Neural Networks with Reproducing Kernel Banach Spaces.Ap- plied and Computational Harmonic Analysis, 62:194–236, January 2023
Francesca Bartolucci, Ernesto De Vito, Lorenzo Rosasco, and Stefano Vigogna. Understanding Neural Networks with Reproducing Kernel Banach Spaces.Ap- plied and Computational Harmonic Analysis, 62:194–236, January 2023
2023
-
[7]
A Duality Framework for Analyzing Random Feature and Two-Layer Neural Networks.The Annals of Statistics, 53(3):1044–1067, June 2025
Hongrui Chen, Jihao Long, and Lei Wu. A Duality Framework for Analyzing Random Feature and Two-Layer Neural Networks.The Annals of Statistics, 53(3):1044–1067, June 2025
2025
-
[8]
Functional Analysis and Partial Differential Equations in Spectral Barron Spaces, July 2025
Mourad Choulli, Shuai Lu, and Hiroshi Takase. Functional Analysis and Partial Differential Equations in Spectral Barron Spaces, July 2025
2025
Show all 72 references
-
[9]
Applied and Numerical Harmonic Analysis
Ole Christensen.An Introduction to Frames and Riesz Bases. Applied and Numerical Harmonic Analysis. Springer International Publishing, Cham, 2016
2016
-
[10]
G. Cybenko. Approximation by Superpositions of a Sigmoidal Function.Math- ematics of Control, Signals and Systems, 2(4):303–314, December 1989
1989
-
[11]
Neural Network Approx- imation.Acta Numerica, 30:327–444, May 2021
Ronald DeVore, Boris Hanin, and Guergana Petrova. Neural Network Approx- imation.Acta Numerica, 30:327–444, May 2021
2021
-
[12]
Ronald A. DeVore. Nonlinear Approximation.Acta Numerica, 7:51–150, Jan- uary 1998. 49
1998
-
[13]
DeVore, Ralph Howard, and Charles Micchelli
Ronald A. DeVore, Ralph Howard, and Charles Micchelli. Optimal Nonlinear Approximation.Manuscripta Mathematica, December 1989
1989
-
[14]
American Mathematical Society, Providence, Rhode Island, De- cember 2000
Javier Duoandikoetxea.Fourier Analysis, volume 29 ofGraduate Studies in Mathematics. American Mathematical Society, Providence, Rhode Island, De- cember 2000
2000
-
[15]
The Barron Space and the Flow-Induced Function Spaces for Neural Network Models.Constructive Approximation, 55(1):369–406, February 2022
Weinan E, Chao Ma, and Lei Wu. The Barron Space and the Flow-Induced Function Spaces for Neural Network Models.Constructive Approximation, 55(1):369–406, February 2022
2022
-
[16]
Friedlander and Mark S
Friedrich G. Friedlander and Mark S. Joshi.Introduction to the Theory of Dis- tributions. Cambridge Univ. Press, Cambridge, 2. ed., transferred to digital printing edition, 2003
2003
-
[17]
London Mathematical Society Lecture Note Series
Alain Grigis and Johannes Sj¨ ostrand.Microlocal Analysis for Differential Op- erators: An Introduction. London Mathematical Society Lecture Note Series. Cambridge University Press, Cambridge, 1994
1994
-
[18]
Guillemin and S
V. Guillemin and S. Sternberg.Geometric Asymptotics, volume 14 ofMathe- matical Surveys and Monographs. American Mathematical Society, Providence, Rhode Island, revised edition, December 1977
1977
-
[19]
Divergence-free Linearized Neural Networks: Integral Representation and Optimal Approximation Rates, 2026
Juncai He, Xinliang Liu, and Zitong Tian. Divergence-free Linearized Neural Networks: Integral Representation and Optimal Approximation Rates, 2026
2026
-
[20]
Springer, New York, NY, 2010
Sigurdur Helgason.Integral Geometry and Radon Transforms. Springer, New York, NY, 2010
2010
-
[21]
Existence and Approximation of Solutions of Differential Equations
Lars H¨ ormander. Existence and Approximation of Solutions of Differential Equations. In Lars H¨ ormander, editor,The Analysis of Linear Partial Dif- ferential Operators II: Differential Operators with Constant Coefficients, pages 3–59. Springer, Berlin, Heidelberg, 2005
2005
-
[22]
Lagrangian Distributions and Fourier Integral Operators
Lars H¨ ormander. Lagrangian Distributions and Fourier Integral Operators. In Lars H¨ ormander, editor,The Analysis of Linear Partial Differential Operators IV: Fourier Integral Operators, pages 3–53. Springer, Berlin, Heidelberg, 2009
2009
-
[23]
Approximation Capabilities of Multilayer Feedforward Networks
Kurt Hornik. Approximation Capabilities of Multilayer Feedforward Networks. Neural Networks, 4(2):251–257, January 1991. 50
1991
-
[24]
Lee K. Jones. A Simple Lemma on Greedy Approximation in Hilbert Space and Convergence Rates for Projection Pursuit Regression and Neural Network Training.The Annals of Statistics, 20(1):608–613, March 1992
1992
-
[25]
Klusowski and Andrew R
Jason M. Klusowski and Andrew R. Barron. Approximation by Combinations of ReLU and Squared ReLU Ridge Functions Withℓ 1 andℓ 0 Controls.IEEE Transactions on Information Theory, 64(12):7649–7656, December 2018
2018
-
[26]
John Wiley & Sons, 2014
Peter D Lax.Functional Analysis. John Wiley & Sons, 2014
2014
-
[27]
Springer, Berlin, Heidelberg, 1991
Michel Ledoux and Michel Talagrand.Probability in Banach Spaces. Springer, Berlin, Heidelberg, 1991
1991
-
[28]
Lin, Allan Pinkus, and Shimon Schocken
Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken. Mul- tilayer Feedforward Networks with a Nonpolynomial Activation Function Can Approximate Any Function.Neural Networks, 6(6):861–867, January 1993
1993
-
[29]
Pereverzev
Yuanyuan Li, Shuai Lu, Peter Math´ e, and Sergei V. Pereverzev. Two-Layer Networks with the ReLU k Activation Function: Barron Spaces and Derivative Approximation.Numerische Mathematik, 156(1):319–344, February 2024
2024
-
[30]
Spectral Barron Space for Deep Neural Network Approximation.SIAM Journal on Mathematics of Data Science, July 2025
Yulei Liao and Pingbing Ming. Spectral Barron Space for Deep Neural Network Approximation.SIAM Journal on Mathematics of Data Science, July 2025
2025
-
[31]
Integral representations of Sobolev spaces via ReLU k activation function and optimal error estimates for linearized networks.Mathematics of Computation, to appear, 2025
Xinliang Liu, Tong Mao, and Jinchao Xu. Integral representations of Sobolev spaces via ReLU k activation function and optimal error estimates for linearized networks.Mathematics of Computation, to appear, 2025. arxiv:2505.00351
2025
-
[32]
Deep Network Approximation for Smooth Functions.SIAM Journal on Mathematical Analysis, 53(5):5465–5506, January 2021
Jianfeng Lu, Zuowei Shen, Haizhao Yang, and Shijun Zhang. Deep Network Approximation for Smooth Functions.SIAM Journal on Mathematical Analysis, 53(5):5465–5506, January 2021
2021
-
[33]
The Radon Transform on Euclidean Space.Communications on Pure and Applied Mathematics, 19(1):49–81, 1966
Donald Ludwig. The Radon Transform on Euclidean Space.Communications on Pure and Applied Mathematics, 19(1):49–81, 1966
1966
-
[34]
Siegel, and Jinchao Xu
Limin Ma, Jonathan W. Siegel, and Jinchao Xu. Uniform Approximation Rates and Metric Entropy of Shallow Neural Networks.Research in the Mathematical Sciences, 9(3):46, July 2022
2022
-
[35]
Y. Makovoz. Random Approximants and Neural Networks.Journal of Approx- imation Theory, 85(1):98–109, April 1996. 51
1996
-
[36]
Siegel, and Jinchao Xu
Tong Mao, Jonathan W. Siegel, and Jinchao Xu. Approximation Rates for Shallow ReLUk Neural Networks on Sobolev Spaces via the Radon Transform. SIAM Journal on Mathematical Analysis, 58(2):1171–1186, April 2026
2026
-
[37]
Solving High-Dimensional PDEs Using Linearized Neural Networks, January 2026
Tong Mao, Jinchao Xu, and Xiaofeng Xu. Solving High-Dimensional PDEs Using Linearized Neural Networks, January 2026
2026
-
[38]
B. Maurey. Type et cotype dans les espaces munis de structures locales incon- ditionnelles.S´ eminaire Analyse fonctionnelle (dit ”Maurey-Schwartz”), pages 1–25, 1973-1974
1973
-
[39]
A New Function Space from Barron Class and Application to Neural Network Approximation.Communications in Computa- tional Physics, 32(5):1361–1400, January 2022
Yan Meng and Pingbing Ming. A New Function Space from Barron Class and Application to Neural Network Approximation.Communications in Computa- tional Physics, 32(5):1361–1400, January 2022
2022
-
[40]
A Func- tion Space View of Bounded Norm Infinite Width ReLU Nets: The Multivari- ate Case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro. A Func- tion Space View of Bounded Norm Infinite Width ReLU Nets: The Multivari- ate Case. InInternational Conference on Learning Representations, September 2019
2019
-
[41]
PhD thesis, University of Wisconsin–Madison, 2022
Rahul Parhi.On Ridge Splines, Neural Networks, and Variational Problems in Radon-Domain BV Spaces. PhD thesis, University of Wisconsin–Madison, 2022
2022
-
[42]
Rahul Parhi and Robert D. Nowak. Banach Space Representer Theorems for Neural Networks and Ridge Splines.J. Mach. Learn. Res., 22(1):43:1960– 43:1999, 2021
1960
-
[43]
Rahul Parhi and Robert D. Nowak. What Kinds of Functions Do Deep Neural Networks Learn? Insights from Variational Spline Theory.SIAM Journal on Mathematics of Data Science, 4(2):464–489, June 2022
2022
-
[44]
Distributional Extension and Invertibility of thek-Plane Transform and Its Dual.SIAM Journal on Mathematical Analysis, 56(4):4662–4686, August 2024
Rahul Parhi and Michael Unser. Distributional Extension and Invertibility of thek-Plane Transform and Its Dual.SIAM Journal on Mathematical Analysis, 56(4):4662–4686, August 2024
2024
-
[45]
Juan C. Peral. L p Estimates for the Wave Equation.Journal of Functional Analysis, 36(1):114–145, March 1980
1980
-
[46]
Petrushev
Pencho P. Petrushev. Approximation by Ridge Functions and Neural Networks. SIAM Journal on Mathematical Analysis, 30(1):155–189, January 1998. 52
1998
-
[47]
Springer, Berlin, Heidelberg, 1985
Allan Pinkus.N-Widths in Approximation Theory. Springer, Berlin, Heidelberg, 1985
1985
-
[48]
Approximation Theory of the MLP Model in Neural Networks
Allan Pinkus. Approximation Theory of the MLP Model in Neural Networks. Acta Numerica, 8:143–195, January 1999
1999
-
[49]
G. Pisier. Remarques Sur Un R´ esultat Non Publi´ e de B. Maurey. InSeminaire analyse fonctionnelle, 1981
1981
-
[50]
An Introduction to X-ray Tomography and Radon Trans- forms
Eric Todd Quinto. An Introduction to X-ray Tomography and Radon Trans- forms. In Gestur ´Olafsson and Eric Quinto, editors,Proceedings of Symposia in Applied Mathematics, volume 63, pages 1–23. American Mathematical Society, Providence, Rhode Island, 2006
2006
-
[51]
Random Features for Large-Scale Kernel Ma- chines
Ali Rahimi and Benjamin Recht. Random Features for Large-Scale Kernel Ma- chines. InAdvances in Neural Information Processing Systems, volume 20. Cur- ran Associates, Inc., 2007
2007
-
[52]
Weighted Sums of Random Kitchen Sinks: Replacing Minimization with Randomization in Learning
Ali Rahimi and Benjamin Recht. Weighted Sums of Random Kitchen Sinks: Replacing Minimization with Randomization in Learning. InAdvances in Neural Information Processing Systems, volume 21. Curran Associates, Inc., 2008
2008
-
[53]
Generalization Properties of Learning with Random Features
Alessandro Rudi and Lorenzo Rosasco. Generalization Properties of Learning with Random Features. InProceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, pages 3218–3228, Red Hook, NY, USA, December 2017. Curran Associates Inc
2017
-
[54]
On the Sharpness of Seeger-Sogge-Stein Orders.Hokkaido Mathematical Journal, 28(2):357–362, February 1999
Michael Ruzhansky. On the Sharpness of Seeger-Sogge-Stein Orders.Hokkaido Mathematical Journal, 28(2):357–362, February 1999
1999
-
[55]
Sogge, and Elias M
Andreas Seeger, Christopher D. Sogge, and Elias M. Stein. Regularity Properties of Fourier Integral Operators.Annals of Mathematics, 134(2):231–251, 1991
1991
-
[56]
Deep Network Approxima- tion Characterized by Number of Neurons.Communications in Computational Physics, 28(5):1768–1811, November 2020
Zuowei Shen, Haizhao Yang, and Shijun Zhang. Deep Network Approxima- tion Characterized by Number of Neurons.Communications in Computational Physics, 28(5):1768–1811, November 2020
2020
-
[57]
Optimal Approximation Rate of ReLU Networks in Terms of Width and Depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, January 2022
Zuowei Shen, Haizhao Yang, and Shijun Zhang. Optimal Approximation Rate of ReLU Networks in Terms of Width and Depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, January 2022. 53
2022
-
[58]
Jonathan W. Siegel. Optimal Approximation Rates for Deep ReLU Neural Networks on Sobolev and Besov Spaces.Journal of Machine Learning Research, 24(357):1–52, 2023
2023
-
[59]
Siegel and Jinchao Xu
Jonathan W. Siegel and Jinchao Xu. High-Order Approximation Rates for Shal- low Neural Networks with Cosine and ReLU k Activation Functions, December 2021
2021
-
[60]
Siegel and Jinchao Xu
Jonathan W. Siegel and Jinchao Xu. Optimal Convergence Rates for the Orthog- onal Greedy Algorithm.IEEE Transactions on Information Theory, 68(5):3354– 3361, May 2022
2022
-
[61]
Siegel and Jinchao Xu
Jonathan W. Siegel and Jinchao Xu. Characterization of the Variation Spaces Corresponding to Shallow Neural Networks.Constructive Approximation, 57(3):1109–1132, June 2023
2023
-
[62]
Siegel and Jinchao Xu
Jonathan W. Siegel and Jinchao Xu. Sharp Bounds on the Approximation Rates, Metric Entropy, andn-Widths of Shallow Neural Networks.Foundations of Computational Mathematics, 24(2):481–537, 2024
2024
-
[63]
Sogge.Fourier Integrals in Classical Analysis
Christopher D. Sogge.Fourier Integrals in Classical Analysis. Cambridge Tracts in Mathematics. Cambridge University Press, Cambridge, 2 edition, 2017
2017
-
[64]
Neural Network with Unbounded Activation Functions Is Universal Approximator.Applied and Computational Harmonic Analysis, 43(2):233–268, September 2017
Sho Sonoda and Noboru Murata. Neural Network with Unbounded Activation Functions Is Universal Approximator.Applied and Computational Harmonic Analysis, 43(2):233–268, September 2017
2017
-
[65]
Stein.Harmonic Analysis: Real-Variable Methods, Orthogonality, and Oscillatory Integrals, volume 43 ofPrinceton Mathematical Series
Elias M. Stein.Harmonic Analysis: Real-Variable Methods, Orthogonality, and Oscillatory Integrals, volume 43 ofPrinceton Mathematical Series. Princeton University Press, Princeton, NJ, 1993
1993
-
[66]
Ridges, Neural Networks, and the Radon Transform.The Jour- nal of Machine Learning Research, 24(1):37:1450–37:1482, January 2023
Michael Unser. Ridges, Neural Networks, and the Radon Transform.The Jour- nal of Machine Learning Research, 24(1):37:1450–37:1482, January 2023
2023
-
[67]
From Kernel Methods to Neural Networks: A Unifying Vari- ational Formulation.Foundations of Computational Mathematics, 24(6):1779– 1818, December 2024
Michael Unser. From Kernel Methods to Neural Networks: A Unifying Vari- ational Formulation.Foundations of Computational Mathematics, 24(6):1779– 1818, December 2024
2024
-
[68]
Warner.Foundations of Differentiable Manifolds and Lie Groups, volume 94 ofGraduate Texts in Mathematics
Frank W. Warner.Foundations of Differentiable Manifolds and Lie Groups, volume 94 ofGraduate Texts in Mathematics. Springer, New York, NY, 1983
1983
-
[69]
Embedding Inequalities for Barron-type Spaces, December 2023
Lei Wu. Embedding Inequalities for Barron-type Spaces, December 2023. 54
2023
-
[70]
Finite Neuron Method and Convergence Analysis.Communications in Computational Physics, 28(5):1707–1745, January 2020
Jinchao Xu. Finite Neuron Method and Convergence Analysis.Communications in Computational Physics, 28(5):1707–1745, January 2020
2020
-
[71]
On the Optimal Approximation of Sobolev and Besov Functions Using Deep ReLU Neural Networks.Applied and Computational Harmonic Analysis, 79:101797, October 2025
Yunfei Yang. On the Optimal Approximation of Sobolev and Besov Functions Using Deep ReLU Neural Networks.Applied and Computational Harmonic Analysis, 79:101797, October 2025
2025
-
[72]
Optimal Rates of Approximation by Shal- low ReLU k Neural Networks and Applications to Nonparametric Regression
Yunfei Yang and Ding-Xuan Zhou. Optimal Rates of Approximation by Shal- low ReLU k Neural Networks and Applications to Nonparametric Regression. Constructive Approximation, 62(2):329–360, October 2025. 55
2025
Reviewed June 25, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.