Pith. sign in

REVIEW 2 major objections 6 minor 63 references

For analytic functions, ReLU network depth drives error down as N to a power of L, not as a symmetric product of N and L.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 10:36 UTC pith:D4C5ASF3

load-bearing objection Clean (N,L) rates for analytic targets that make depth strictly more powerful than width, with nearly matching lower bounds and solid constructions. the 2 major comments →

arxiv 2607.10589 v1 pith:D4C5ASF3 submitted 2026-07-12 stat.ML cs.ITcs.LGcs.NAmath.ITmath.NA

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width

classification stat.ML cs.ITcs.LGcs.NAmath.ITmath.NA MSC 68T0741A2562G08
keywords ReLU networksanalytic functionsdepth-width trade-offLegendre approximationnonparametric regressionBernstein polyellipse(N,L)-characterization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most neural-network approximation theory either counts total parameters or treats width N and depth L as interchangeable. For functions with only finite smoothness that interchange is natural: error typically scales like a product of negative powers of N and L. Analytic functions are infinitely smooth and extend holomorphically into a complex neighborhood of the cube. This paper proves that, once that stronger regularity is present, depth becomes strictly more powerful than width. Under the joint (N,L) characterization the approximation error can be made as small as N raised to a negative power that itself grows with a positive power of L. In the regime where width is roughly a power of depth the exponent on L reaches 1, producing an essentially exponential-in-depth rate. The same rates are nearly matched by a matching lower bound when width grows like a high power of depth, and they translate into near-minimax nonparametric regression rates that are essentially 1/n up to logs. The technical engine is a careful construction that first approximates powers and multiplications by shallow ReLU gadgets and then assembles multivariate Legendre expansions without letting intermediate errors explode.

Core claim

For any f belonging to the analytic class A(rho,M) and for N,L large enough with L to a power kappa-plus-alpha less than or equal to N less than or equal to exp(L to beta), there exists a ReLU network of width proportional to N and depth proportional to L (or L times (log L) squared when kappa equals d) whose uniform error on the cube is at most C times N to the power minus C L to the power tau(kappa), where tau is the piecewise-linear function of kappa given in Theorem 1. In particular, depth dominates width.

What carries the argument

Refined ReLU constructions that approximate univariate powers (binary-tree multiplication), multivariate products, and Legendre polynomials, then assemble a truncated multivariate Legendre series of the analytic target while controlling the trade-off between polynomial degree and network size.

Load-bearing premise

The target function must extend holomorphically to a fixed Bernstein polyellipse whose size is known in advance and stay bounded there; if the analyticity radius is smaller or only finite smoothness holds, the claimed depth-driven rates disappear.

What would settle it

Compute the best uniform approximation error of a concrete analytic function (for example a scaled product of cosines) by exhaustive search or mixed-integer programming over all ReLU networks of width N and depth L in a regime where N is only polynomial in L; if the observed error decays no faster than any fixed power of 1/(N L), the claimed depth advantage is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper derives upper bounds on the uniform approximation of analytic functions f in the class A(rho, M) (holomorphic extensions to a fixed Bernstein polyellipse with bounded L^infty norm) by ReLU networks of width O(N) and depth O(L). Under the regime L^{kappa+alpha} <= N <= exp(L^beta), the rate is O(N^{-C L^{tau(kappa)}}) with a piecewise-linear exponent tau(kappa) that is strictly larger when depth is relatively larger; in particular tau=1 when N scales like L^d. This shows depth is more powerful than width for infinite-smoothness targets, in contrast to the symmetric N^{-2s/d} L^{-2s/d} rates known for finite-smoothness classes under the same (N,L) characterization. The proofs rest on multivariate Legendre expansions, refined binary-tree ReLU constructions for powers and multiplications (Lemmas 6-7), two distinct univariate Legendre approximators (Lemmas 8 and 11), and a squashing technique for linear combinations (Lemmas 10 and 12). A matching-order lower bound (Theorem 2) is given via convexity and the isodiametric inequality, and the approximation theory is applied to obtain nearly-minimax rates for nonparametric regression with analytic targets (Theorems 3-4).

Significance. The work fills a clear gap: prior (N,L)-characterization results treated only finite smoothness, while existing analytic-function approximation results used single-parameter (depth or total size) characterizations. The constructive intermediate networks for powers, multiplications and Legendre polynomials are of independent interest and improve depth dependence relative to earlier constructions. The depth-dominance message is cleanly supported by both the upper bounds and the lower bound that rules out an e^{-N L}-type rate. The nonparametric application yields eO(n^{-1}) excess risk with only sub-polynomial network size, which is a sharp contrast to the polynomial-size networks needed for finite-smoothness classes. The manuscript is fully constructive, cites the classical analytic-approximation and covering-number ingredients correctly, and explicitly flags the remaining open problems (improving the piecewise tau, closing the minor gap on the L-N plane, optimality for kappa < d).

major comments (2)
  1. Theorem 1 and the subsequent comparison with Theorem 2: the upper bound is N^{-C L^{tau(kappa)}} while the lower bound (after the width/depth rescaling stated in Section 1.2) is only Omega(N^{-C L}) (or Omega(N^{-C L (log L)^2}) when kappa=d). The authors correctly note near-optimality only for kappa=d (ignoring logs). For kappa < d the gap between tau(kappa) and 1 is left open; a short remark clarifying whether the lower-bound technique can be sharpened, or whether the upper-bound constructions are believed to be rate-optimal, would strengthen the claim that depth is 'more critical'.
  2. Section 2.4-2.5 (Lemmas 10 and 12): the choice of the free parameters |Lambda_epsilon| ~ L^d (log N)^d, k ~ log N, p ~ L^lambda (log N)^d, q ~ sqrt(N) produces the claimed rates, but the dependence of the hidden constants C(d,rho,kappa,alpha) on alpha (the arbitrarily small exponent in L^{kappa+alpha} <= N) is not tracked. Because alpha appears in the final exponent of N, a one-sentence statement that the constant remains positive for every fixed alpha > 0 (or alpha=0 when kappa=0) would make the result fully self-contained.
minor comments (6)
  1. Figure 1a: the vertical axis is labeled tau(kappa) but the caption says 'An illustration of kappa(tau)'. Swap or correct the caption.
  2. Lemma 1: the bound sum |c_ell^{(nu)}| <= 4^nu is cited from [36, Prop. 4.6]; a one-line sketch or reference to the generating-function argument would help readers who do not have that paper at hand.
  3. Page 7, line after Theorem 2: 'replacing N by C(d,rho)N and L by C(d,rho,beta)L … yields a lower bound of Omega(N^{-C(d,rho,beta)L})'. The constant in the exponent should be written C'(d,rho,beta) to avoid confusion with the C appearing in the upper bound.
  4. Section 3, Theorem 3(III): the depth bound contains an extra (log log n)^2 factor that is absorbed into the eO notation of the excess-risk statement; stating this absorption explicitly would improve readability.
  5. References: [59] and [27] appear as arXiv preprints dated 2025-2026; if they have since been published, update the bibliographic data.
  6. Notation: the same symbol L is used both for network depth and for Legendre polynomials; a brief reminder at the first occurrence of L_nu would reduce momentary confusion.

Circularity Check

0 steps flagged

No significant circularity: rates follow from explicit ReLU constructions of powers/multiplications/Legendre polynomials plus classical holomorphic approximation, with no fitted parameters or load-bearing self-definitions.

full rationale

The central claims (Theorem 1 upper bounds under the (N,L) regime with piecewise τ(κ), Theorem 2 lower bound, and the nonparametric rates in Theorems 3–4) are obtained by fully constructive arguments: binary-tree ReLU networks for powers (Lemma 6) and multivariate multiplication (Lemma 7), inductive constructions for univariate and multivariate Legendre polynomials (Lemmas 8–12), linear combination via squashing, and the classical Legendre truncation error for functions holomorphic on a Bernstein polyellipse (Proposition 3, citing external results such as Opschoor–Schwab–Zech). The lower bound uses only the piecewise-linear partition property of ReLU nets plus the isodiametric inequality. No parameter is fitted to data and then re-used as a “prediction”; no uniqueness theorem is imported from the authors’ own prior work to force the architecture; and the few citations to earlier (N,L) papers (e.g., Shen et al.) or analytic-approximation papers serve only as black-box lemmas whose statements are independent of the present rates. The derivation is therefore self-contained against external mathematical benchmarks and exhibits none of the six circularity patterns.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The paper rests on standard real/complex analysis (holomorphic extension to Bernstein ellipses, Legendre expansions, isodiametric inequality) plus the classical ReLU multiplication lemma of Lu et al. No free parameters are fitted; the only domain assumptions are membership in A(rho,M) and the scaling relation between N and L. No new physical or mathematical entities are postulated.

axioms (4)
  • domain assumption Any real-analytic f on [-1,1]^d extends holomorphically to a Bernstein polyellipse E_rho with ||f||_infty(E_rho) <= M (definition of A(rho,M)).
    Standing assumption of Section 1.2; all rates are stated only for this class.
  • standard math Multivariate Legendre polynomials form an orthonormal basis of L2(mu_d) and their coefficients decay as product rho_i^{-nu_i} (Lemmas 2-3, Prop. 3).
    Classical orthogonal-polynomial theory used as black box.
  • standard math There exists a width-10N depth-L ReLU network that multiplies two scalars on a compact interval with error O(N^{-L}) (Lemma 5, cited from Lu et al. 2021).
    Imported building block for all subsequent constructions.
  • standard math A ReLU network of width N and depth L partitions R^d into at most C N^d L convex polytopes on each of which it is affine (Lemma 13).
    Used only for the lower bound (Thm 2).

pith-pipeline@v1.1.0-grok45 · 45725 in / 2521 out tokens · 28111 ms · 2026-07-14T10:36:43.134784+00:00 · methodology

0 comments
read the original abstract

In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width parameter $N$ and the depth parameter $L$, thereby granting greater architectural flexibility. Existing works using the $(N,L)$-characterization focus on function classes with finite smoothness $s$, establishing a typical approximation rate of $\mathcal{O}\left(N^{-2s/d}L^{-2s/d}\right)$ with $d$ denoting the input dimension, which indicates that network depth and width play symmetric roles for these classes. In contrast, this paper establishes upper bounds for the approximation of analytic functions, which possess infinite smoothness, via ReLU networks under the $(N,L)$-characterization. Specifically, we derive approximation rates of $\mathcal{O}\left(N^{-C L^{\tau}}\right)$, where $C>0$ is some constant and $\tau>0$ is a parameter influenced by the relation between $L$ and $N$. In particular, $\tau=1$ if $N$ scales roughly as $L^d$. Our findings reveal that depth plays a more critical role than width in the context of analytic function approximation. The main technical difficulty of obtaining such upper bounds lies in the trade-off between the smoothness parameters and the approximation accuracy. To overcome this difficulty, we employ refined constructions of several ReLU networks to approximate power functions, multivariate multiplication, and polynomials, which may be of independent interest.

Figures

Figures reproduced from arXiv: 2607.10589 by Defeng Sun, Yang Wang, Yanming Lai.

Figure 1
Figure 1. Figure 1: Visualization of Theorem 1 number yields a network with depth Oe(L) and Oe [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: An illustration of the proof of Lemma 8. [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: An illustration of the proof of Lemma 11. [PITH_FULL_IMAGE:figures/full_fig_p026_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: An illustration of the proof of Lemma 12. [PITH_FULL_IMAGE:figures/full_fig_p029_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

63 extracted references · 3 linked inside Pith

  1. [1]

    SIAM, 2022

    Ben Adcock, Simone Brugiapaglia, and Clayton G Webster.Sparse polynomial ap- proximation of high-dimensional functions, volume 25. SIAM, 2022

  2. [2]

    cambridge university press, 2009

    Martin Anthony and Peter L Bartlett.Neural network learning: Theoretical foun- dations. cambridge university press, 2009

  3. [3]

    Universal approximation bounds for superpositions of a sigmoidal function.IEEE Transactions on Information theory, 39(3):930–945, 1993

    Andrew R Barron. Universal approximation bounds for superpositions of a sigmoidal function.IEEE Transactions on Information theory, 39(3):930–945, 1993

  4. [4]

    Nearly- tight vc-dimension and pseudodimension bounds for piecewise linear neural networks

    Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian. Nearly- tight vc-dimension and pseudodimension bounds for piecewise linear neural networks. Journal of Machine Learning Research, 20(63):1–17, 2019

  5. [5]

    Shallow and deep networks are near-optimal approximators of korobov functions

    Moise Blanchard and Mohammed Amine Bennouna. Shallow and deep networks are near-optimal approximators of korobov functions. InInternational conference on learning representations, 2021

  6. [6]

    Optimal approximation with sparsely connected deep neural networks.SIAM Journal on Mathematics of Data Science, 1(1):8–45, 2019

    Helmut Bolcskei, Philipp Grohs, Gitta Kutyniok, and Philipp Petersen. Optimal approximation with sparsely connected deep neural networks.SIAM Journal on Mathematics of Data Science, 1(1):8–45, 2019

  7. [7]

    Minshuo Chen, Haoming Jiang, Wenjing Liao, and Tuo Zhao. Nonparametric re- gression on low-dimensional manifolds using deep relu networks: Function approxi- mation and statistical recovery.Information and Inference: A Journal of the IMA, 11(4):1203–1253, 2022

  8. [8]

    Breaking the curse of di- mensionality in sparse polynomial approximation of parametric pdes.Journal de Math´ ematiques Pures et Appliqu´ ees, 103(2):400–428, 2015

    Abdellah Chkifa, Albert Cohen, and Christoph Schwab. Breaking the curse of di- mensionality in sparse polynomial approximation of parametric pdes.Journal de Math´ ematiques Pures et Appliqu´ ees, 103(2):400–428, 2015

  9. [9]

    Analytic regularity and poly- nomial approximation of parametric and stochastic elliptic pde’s.Analysis and Ap- plications, 9(01):11–47, 2011

    Albert Cohen, Ronald Devore, and Christoph Schwab. Analytic regularity and poly- nomial approximation of parametric and stochastic elliptic pde’s.Analysis and Ap- plications, 9(01):11–47, 2011

  10. [10]

    Approximation by superpositions of a sigmoidal function.Mathe- matics of control, signals and systems, 2(4):303–314, 1989

    George Cybenko. Approximation by superpositions of a sigmoidal function.Mathe- matics of control, signals and systems, 2(4):303–314, 1989

  11. [11]

    Neural network approximation

    Ronald DeVore, Boris Hanin, and Guergana Petrova. Neural network approximation. Acta Numerica, 30:327–444, 2021

  12. [12]

    Exponential convergence of the deep neural network approximation for analytic functions.Science China Mathematics, 61(10):1733–1740, 2018

    Weinan E and Qingcan Wang. Exponential convergence of the deep neural network approximation for analytic functions.Science China Mathematics, 61(10):1733–1740, 2018

  13. [13]

    Chapman and Hall/CRC, 2015

    Lawrence Craig Evans and Ronald F Gariepy.Measure Theory and Fine Properties of Functions. Chapman and Hall/CRC, 2015

  14. [14]

    Error bounds for approxima- tions with deep relu neural networks in w s, p norms.Analysis and Applications, 18(05):803–859, 2020

    Ingo G¨ uhring, Gitta Kutyniok, and Philipp Petersen. Error bounds for approxima- tions with deep relu neural networks in w s, p norms.Analysis and Applications, 18(05):803–859, 2020. 43

  15. [15]

    Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Networks, 134:107–130, 2021

    Ingo G¨ uhring and Mones Raslan. Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Networks, 134:107–130, 2021

  16. [16]

    Decision theoretic generalizations of the pac model for neural net and other learning applications.Information and computation, 100(1):78–150, 1992

    David Haussler. Decision theoretic generalizations of the pac model for neural net and other learning applications.Information and computation, 100(1):78–150, 1992

  17. [17]

    Sphere packing numbers for subsets of the boolean n-cube with bounded vapnik-chervonenkis dimension.Journal of Combinatorial Theory, Series A, 69(2):217–232, 1995

    David Haussler. Sphere packing numbers for subsets of the boolean n-cube with bounded vapnik-chervonenkis dimension.Journal of Combinatorial Theory, Series A, 69(2):217–232, 1995

  18. [18]

    Simultaneous neural network approximation for smooth functions.Neural Networks, 154:152–164, 2022

    Sean Hon and Haizhao Yang. Simultaneous neural network approximation for smooth functions.Neural Networks, 154:152–164, 2022

  19. [19]

    Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257, 1991

    Kurt Hornik. Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257, 1991

  20. [20]

    On estimation of analytic functions.Studia Sci Math Hungarica, 34:191–210, 1998

    Ildar Ibragimov. On estimation of analytic functions.Studia Sci Math Hungarica, 34:191–210, 1998

  21. [21]

    Springer Science & Business Media, 2013

    Ildar Abdulovich Ibragimov and Rafail Zalmanovich Has’ Minskii.Statistical esti- mation: asymptotic theory. Springer Science & Business Media, 2013

  22. [22]

    Yuling Jiao, Yanming Lai, Xiliang Lu, Fengru Wang, Jerry Zhijian Yang, and Yuanyuan Yang. Deep neural networks with relu-sine-exponential activations break curse of dimensionality in approximation on h¨ older class.SIAM Journal on Mathe- matical Analysis, 55(4):3635–3649, 2023

  23. [23]

    Deep nonparametric regression on approximate manifolds: Nonasymptotic error bounds with polynomial prefactors.The Annals of Statistics, 51(2):691–716, 2023

    Yuling Jiao, Guohao Shen, Yuanyuan Lin, and Jian Huang. Deep nonparametric regression on approximate manifolds: Nonasymptotic error bounds with polynomial prefactors.The Annals of Statistics, 51(2):691–716, 2023

  24. [24]

    Approximation bounds for norm con- strained neural networks with applications to regression and gans.Applied and Computational Harmonic Analysis, 65:249–278, 2023

    Yuling Jiao, Yang Wang, and Yunfei Yang. Approximation bounds for norm con- strained neural networks with applications to regression and gans.Applied and Computational Harmonic Analysis, 65:249–278, 2023

  25. [25]

    On the rate of convergence of fully connected deep neural network regression estimates.The Annals of Statistics, 49(4):2231–2249, 2021

    Michael Kohler and Sophie Langer. On the rate of convergence of fully connected deep neural network regression estimates.The Annals of Statistics, 49(4):2231–2249, 2021

  26. [26]

    Multilayer feedforward networks with a nonpolynomial activation function can approximate any function.Neural networks, 6(6):861–867, 1993

    Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function.Neural networks, 6(6):861–867, 1993

  27. [27]

    Some super-approximation rates of relu neural net- works for korobov functions.arXiv preprint arXiv:2507.10345, 2025

    Yuwen Li and Guozhi Zhang. Some super-approximation rates of relu neural net- works for korobov functions.arXiv preprint arXiv:2507.10345, 2025

  28. [28]

    Shiyu Liang and R. Srikant. Why deep neural networks for function approximation? InInternational Conference on Learning Representations, 2017

  29. [29]

    Deep network approxi- mation for smooth functions.SIAM Journal on Mathematical Analysis, 53(5):5465– 5506, 2021

    Jianfeng Lu, Zuowei Shen, Haizhao Yang, and Shijun Zhang. Deep network approxi- mation for smooth functions.SIAM Journal on Mathematical Analysis, 53(5):5465– 5506, 2021. 44

  30. [30]

    The ex- pressive power of neural networks: A view from the width.Advances in neural information processing systems, 30, 2017

    Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang. The ex- pressive power of neural networks: A view from the width.Advances in neural information processing systems, 30, 2017

  31. [31]

    Rates of approximation by relu shallow neural networks.Journal of Complexity, 79:101784, 2023

    Tong Mao and Ding-Xuan Zhou. Rates of approximation by relu shallow neural networks.Journal of Complexity, 79:101784, 2023

  32. [32]

    Neural networks for optimal approximation of smooth and analytic functions.Neural computation, 8(1):164–177, 1996

    Hrushikesh N Mhaskar. Neural networks for optimal approximation of smooth and analytic functions.Neural computation, 8(1):164–177, 1996

  33. [33]

    Approximation properties of a multilayered feedforward artificial neural network.Advances in Computational Mathematics, 1(1):61–80, 1993

    Hrushikesh Narhar Mhaskar. Approximation properties of a multilayered feedforward artificial neural network.Advances in Computational Mathematics, 1(1):61–80, 1993

  34. [34]

    New error bounds for deep relu networks using sparse grids.SIAM Journal on Mathematics of Data Science, 1(1):78–92, 2019

    Hadrien Montanelli and Qiang Du. New error bounds for deep relu networks using sparse grids.SIAM Journal on Mathematics of Data Science, 1(1):78–92, 2019

  35. [35]

    Adaptive approximation and generalization of deep neural network with intrinsic dimensionality.Journal of Machine Learning Research, 21(174):1–38, 2020

    Ryumei Nakada and Masaaki Imaizumi. Adaptive approximation and generalization of deep neural network with intrinsic dimensionality.Journal of Machine Learning Research, 21(174):1–38, 2020

  36. [36]

    Deep relu networks and high-order finite element methods.Analysis and Applications, 18(05):715–770, 2020

    Joost AA Opschoor, Philipp C Petersen, and Christoph Schwab. Deep relu networks and high-order finite element methods.Analysis and Applications, 18(05):715–770, 2020

  37. [37]

    Exponential relu dnn expression of holomorphic maps in high dimension.Constructive Approximation, 55(1):537–582, 2022

    Joost AA Opschoor, Ch Schwab, and Jakob Zech. Exponential relu dnn expression of holomorphic maps in high dimension.Constructive Approximation, 55(1):537–582, 2022

  38. [38]

    Optimal approximation of piecewise smooth functions using deep relu neural networks.Neural Networks, 108:296–330, 2018

    Philipp Petersen and Felix Voigtlaender. Optimal approximation of piecewise smooth functions using deep relu neural networks.Neural Networks, 108:296–330, 2018

  39. [39]

    Approximation theory of the mlp model in neural networks.Acta numerica, 8:143–195, 1999

    Allan Pinkus. Approximation theory of the mlp model in neural networks.Acta numerica, 8:143–195, 1999

  40. [40]

    On the expressive power of deep neural networks

    Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl- Dickstein. On the expressive power of deep neural networks. Ininternational con- ference on machine learning, pages 2847–2854. PMLR, 2017

  41. [41]

    Nonparametric regression using deep neural networks with relu activation function.The Annals of Statistics, 48(4):1875, 2020

    Johannes Schmidt-Hieber. Nonparametric regression using deep neural networks with relu activation function.The Annals of Statistics, 48(4):1875, 2020

  42. [42]

    Bounding and counting linear regions of deep neural networks

    Thiago Serra, Christian Tjandraatmadja, and Srikumar Ramalingam. Bounding and counting linear regions of deep neural networks. InInternational conference on machine learning, pages 4558–4566. PMLR, 2018

  43. [43]

    Expressivity of shallow and deep neural networks for polynomial ap- proximation.arXiv preprint arXiv:2303.03544, 2023

    Itai Shapira. Expressivity of shallow and deep neural networks for polynomial ap- proximation.arXiv preprint arXiv:2303.03544, 2023

  44. [44]

    Deep network approximation characterized by number of neurons.Communications in Computational Physics, 28(5):1768–1811, 2020

    Zuowei Shen, Haizhao Yang, and Shijun Zhang. Deep network approximation characterized by number of neurons.Communications in Computational Physics, 28(5):1768–1811, 2020. 45

  45. [45]

    Optimal approximation rate of relu networks in terms of width and depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, 2022

    Zuowei Shen, Haizhao Yang, and Shijun Zhang. Optimal approximation rate of relu networks in terms of width and depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, 2022

  46. [46]

    Optimal approximation rates for deep relu neural networks on sobolev and besov spaces.Journal of Machine Learning Research, 24(357):1–52, 2023

    Jonathan W Siegel. Optimal approximation rates for deep relu neural networks on sobolev and besov spaces.Journal of Machine Learning Research, 24(357):1–52, 2023

  47. [47]

    Approximation rates for neural networks with general activation functions.Neural Networks, 128:313–321, 2020

    Jonathan W Siegel and Jinchao Xu. Approximation rates for neural networks with general activation functions.Neural Networks, 128:313–321, 2020

  48. [48]

    High-order approximation rates for shallow neu- ral networks with cosine and reluk activation functions.Applied and Computational Harmonic Analysis, 58:1–26, 2022

    Jonathan W Siegel and Jinchao Xu. High-order approximation rates for shallow neu- ral networks with cosine and reluk activation functions.Applied and Computational Harmonic Analysis, 58:1–26, 2022

  49. [49]

    Optimal global rates of convergence for nonparametric regression

    Charles J Stone. Optimal global rates of convergence for nonparametric regression. The annals of statistics, pages 1040–1053, 1982

  50. [50]

    Adaptivity of deep relu network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality

    Taiji Suzuki. Adaptivity of deep relu network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality. InInternational Conference on Learning Representations, 2019

  51. [51]

    Deep learning is adaptive to intrinsic dimensional- ity of model smoothness in anisotropic besov space.Advances in Neural Information Processing Systems, 34:3609–3621, 2021

    Taiji Suzuki and Atsushi Nitanda. Deep learning is adaptive to intrinsic dimensional- ity of model smoothness in anisotropic besov space.Advances in Neural Information Processing Systems, 34:3609–3621, 2021

  52. [52]

    Representation benefits of deep feedforward networks.arXiv preprint arXiv:1509.08101, 2015

    Matus Telgarsky. Representation benefits of deep feedforward networks.arXiv preprint arXiv:1509.08101, 2015

  53. [53]

    Springer series in statistics

    Alexandre B Tsybakov.Introduction to Nonparametric Estimation. Springer series in statistics. Springer, Dordrecht, 2009

  54. [54]

    Cambridge university press, 2019

    Martin J Wainwright.High-dimensional statistics: A non-asymptotic viewpoint, volume 48. Cambridge university press, 2019

  55. [55]

    Deep neural networks with general activations: Super- convergence in sobolev norms.arXiv preprint arXiv:2508.05141, 2025

    Yahong Yang and Juncai He. Deep neural networks with general activations: Super- convergence in sobolev norms.arXiv preprint arXiv:2508.05141, 2025

  56. [56]

    Near-optimal deep neural network approximation for korobov functions with respect to lp and h1 norms.Neural Networks, 180:106702, 2024

    Yahong Yang and Yulong Lu. Near-optimal deep neural network approximation for korobov functions with respect to lp and h1 norms.Neural Networks, 180:106702, 2024

  57. [57]

    Nearly optimal vc-dimension and pseudo-dimension bounds for deep neural network derivatives.Advances in Neural Information Processing Systems, 36:21721–21756, 2023

    Yahong Yang, Haizhao Yang, and Yang Xiang. Nearly optimal vc-dimension and pseudo-dimension bounds for deep neural network derivatives.Advances in Neural Information Processing Systems, 36:21721–21756, 2023

  58. [58]

    On the optimal approximation of sobolev and besov functions using deep relu neural networks.Applied and Computational Harmonic Analysis, page 101797, 2025

    Yunfei Yang. On the optimal approximation of sobolev and besov functions using deep relu neural networks.Applied and Computational Harmonic Analysis, page 101797, 2025

  59. [59]

    Approximation and learning of anisotropic and mixed smooth functions by deep relu neural networks.arXiv preprint arXiv:2605.31152, 2026

    Yunfei Yang and Jun Fan. Approximation and learning of anisotropic and mixed smooth functions by deep relu neural networks.arXiv preprint arXiv:2605.31152, 2026. 46

  60. [60]

    Optimal rates of approximation by shallow relu k neural networks and applications to nonparametric regression.Constructive Approximation, 62(2):329–360, 2025

    Yunfei Yang and Ding-Xuan Zhou. Optimal rates of approximation by shallow relu k neural networks and applications to nonparametric regression.Constructive Approximation, 62(2):329–360, 2025

  61. [61]

    Error bounds for approximations with deep relu networks.Neural networks, 94:103–114, 2017

    Dmitry Yarotsky. Error bounds for approximations with deep relu networks.Neural networks, 94:103–114, 2017

  62. [62]

    Optimal approximation of continuous functions by very deep relu networks

    Dmitry Yarotsky. Optimal approximation of continuous functions by very deep relu networks. InConference on learning theory, pages 639–649. PMLR, 2018

  63. [63]

    Deep network approximation: Be- yond relu to diverse activation functions.Journal of Machine Learning Research, 25(35):1–39, 2024

    Shijun Zhang, Jianfeng Lu, and Hongkai Zhao. Deep network approximation: Be- yond relu to diverse activation functions.Journal of Machine Learning Research, 25(35):1–39, 2024. 47