Pith. sign in

REVIEW 3 major objections 6 minor 66 references

Approximation Rates in Fr\'echet Metrics: Barron Spaces, Paley-Wiener Spaces, and Fourier Multipliers

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Shallow networks can approximate Fourier symbols of linear operators in Fréchet metrics to any prescribed accuracy, with explicit sufficient widths, and the Barron-bandlimited class attains the dimension-independent $N^{-1/2}$ Monte Carlo…

desk verdict A clean Fréchet-width theorem that is currently oversold by a broken bandlimited example and an overclaimed novelty statement. read the letter →

arxiv 2501.04023 v2 pith:IHH7BNAB submitted 2024-12-27 math.NA cs.ITcs.LGcs.NAmath.ITstat.ML

classification math.NAcs.ITcs.LGcs.NAmath.ITstat.ML MSC 41A2541A4641A6546E1068T0568T07
keywords NeuralnetworksApproximationratesOperatorlearningSymbolBarronspacesFréchetPaley-WienerFouriermultipliers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies how well shallow neural networks can approximate the Fourier symbols of linear differential operators, where the error is measured not by a single norm but by a whole Fréchet metric built from a sequence of semi-norms. The main theorems give sufficient conditions, in closed form, on the network width $N$ needed to drive the Fréchet error below any prescribed $\varepsilon$: one theorem requires approximation rates in every semi-norm under a mild monotonic-growth condition, and a second theorem requires only the rate for the first semi-norm under a stronger bounded-growth condition. Applied to exponentially weighted spectral Barron spaces and Gelfand-Shilov spaces, the first theorem yields explicit width formulas; applied to a new class of Barron-bandlimited functions, the second yields a dimension-independent width formula with the Monte Carlo rate $N^{-1/2}$. The upshot is that infinite-dimensional operator approximation can be reduced to explicit finite-network-size guarantees, with the caveat that the bounded-growth condition is narrow and fails for the natural exponential-spectral-Barron example.

What carries the argument

The central object is the Fréchet metric $d_V(f)=\sum_{\ell\in\mathbb{Z}_+} 2^{-\ell} p_\ell(f)/(1+p_\ell(f))$ induced by a separating sequence of semi-norms $\{p_\ell\}$, which lets the authors measure symbol approximation in a topology strictly finer-grained than any single Banach norm. The proof machinery splits the infinite sum in the metric at some level $\ell$: the tail is bounded by $2^{-\ell}$ by direct geometric summation, while the head is bounded through the monotonic-growth or bounded-growth condition together with the monotonicity of $t\mapsto t/(1+t)$ and the assumed rate $r_\ell$. Inverting the rate gives the closed-form width $N_{\ell_\varepsilon}$ for any target accuracy $\varepsilon$. The bounded-growth condition $p_\ell(f)\le M_\ell p_0(f)$ is the specific bridge that the second theorem uses to transfer an approximation rate known only for $p_0$ to the whole Fréchet metric.

What would settle it

Compute the best $N$-term approximation error in the $L^2$ semi-norm for a concrete Barron-bandlimited function, for instance a function whose Fourier transform is a truncated Gaussian on $[-\Omega,\Omega]^d$, and check whether the error decays as $N^{-1/2}$ with the implied constant independent of dimension; if the decay is slower, Proposition 4.8 would be contradicted. For Theorem 4.9, evaluate the full Fréchet metric $d_V(f-f_N)$ for the width given by the closed-form formula and test whether it indeed falls below $\varepsilon$; a single violation would falsify the suffiency claim.

Watch

Extended reading notes

Core claim

At the center of the paper are two general approximation theorems. Theorem 3.1 assumes the monotonic-growth condition $p_k(f) \le M_\ell p_\ell(f)$ for $k \le \ell$ and, for every semi-norm, a decreasing bijective rate $r_\ell$ with $p_\ell(f-f_N) \le C_f r_\ell(N)$; it then proves that choosing $N \ge r_{\ell_\varepsilon}^{-1}\bigl(\min\{r_{\ell_\varepsilon}(1), 2^{-\ell_\varepsilon}/(C_f M_{\ell_\varepsilon})\bigr)\bigr)$, where $\ell_\varepsilon = \lceil -\log_2\varepsilon\rceil + 1$, guarantees $d_V(f-f_N)<\varepsilon$. Theorem 4.1 replaces the monotonic condition by the bounded-growth condition $p_\ell(f)\le M_\ell p_0(f)$ and uses only the rate for $p_0$, with the sufficient width $N \ge r^{-1}\bigl(\min\{r(1), 2^{-\ell_\varepsilon}/(C_f M_{\ell_\varepsilon})\bigr)\bigr)$. The authors then show that the exponential spectral Barron space satisfies the monotonic condition with Sobolev semi-norms, producing a width of order $\bigl(\tfrac{1}{c_{\ell_\varepsilon}}\ln(2^{\ell_\varepsilon}C_{\ell_\varepsilon}\|f\|_{B_{\beta,c}})\bigr)^{d/\beta}$, but that it violates the bounded-growth condition: cosine units are not in the space for $\beta\ge 1$, and derivatives are unbounded for $\beta<1$. For Barron-bandlimited functions $B_1^*(\Omega)$, the bounded-growth condition holds with $M_\ell = \langle\Omega\rangle^\ell$, and a dimension-independent width $N \ge \bigl(2^{-\ell_\varepsilon}/(\|f\|_{B_1^*(\Omega)}\langle\Omega\rangle^{\ell_\varepsilon})\bigr)^{-2}$ guarantees the Fréchet error bound.

Load-bearing premise

The load-bearing premise of the second main theorem is that each derivative-level semi-norm of the target function is bounded by a fixed multiple of the semi-norm at level zero; the paper shows that this fails for the natural exponential spectral Barron space, so the theorem only applies to classes specially made to satisfy it, such as the Barron-bandlimited class.

Editorial extensions

If this is right

  • If the two main theorems are correct, then every symbol class with known rate functions in each semi-norm, such as the exponential spectral Barron space and the Gelfand-Shilov spaces considered here, has an explicit shallow-network width for a preassigned Fréchet accuracy.
  • If the bounded-growth condition holds for a class, a single $L^2$ (or first semi-norm) approximation rate suffices to control the entire Fréchet metric, so existing Monte Carlo rates transfer to the infinite-dimensional symbol-approximation setting.
  • For the Barron-bandlimited class, the sufficient width formula gives approximation in all Sobolev semi-norms simultaneously with the dimension-independent rate $N^{-1/2}$.
  • The counterexamples for the exponential spectral Barron space show that the bounded-growth condition is not automatically satisfied by natural classes, so applications of the second theorem must verify condition (4.1) explicitly.
  • The embedding results connect Barron spaces to the multiplier class $S^{1,\ell}(U)$ and to Gelfand-Shilov spaces, so the approximation guarantees apply to standard classes of pseudodifferential symbols.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the width formula in Theorem 4.9 still depends on the bandwidth through $M_{\ell_\varepsilon}=\langle\Omega\rangle^{\ell_\varepsilon}$; for families of symbols whose bandwidth grows with dimension or target smoothness, the sufficient width inherits that growth, so the "no curse of dimensionality" claim applies to fixed bandwidth at fixed accuracy.
  • Editorial inference: the paper approximates symbols, not solution operators end-to-end; in an operator-learning pipeline, the Fréchet symbol error would need to be combined with a stability estimate for the underlying PDE to guarantee control of the operator's output error, which the paper leaves implicit.
  • Editorial inference: Definition 4.7 writes the class condition with $B^1$ but the norm with $B^s$; a natural clarification is whether the rate in Proposition 4.8 and the width in Theorem 4.9 require $s=1$ or hold uniformly over $s>1$.
  • Editorial inference: a natural next step would be to characterize which weighted symbol classes satisfy the bounded-growth condition while remaining rich enough to contain non-trivial approximation classes, since the self-weighted Gevrey class considered in the appendix turns out to be trivial or intractable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper studies neural-network approximation of symbols/functions measured in a Fréchet metric induced by a separating sequence of semi-norms. The two main results are meta-theorems: Theorem 3.1 shows that if approximation rates are known for every semi-norm and the seminorms satisfy a monotonic-growth condition (3.1), then an explicit width N suffices to make the Fréchet error dV(f-f_N) < ε; Theorem 4.1 shows that under the stronger bounded-growth condition (4.1), knowledge of the rate in only the first semi-norm suffices, again with an explicit width. The paper applies the first result to the exponential spectral Barron space with Sobolev semi-norms and to Gelfand-Shilov spaces, and the second result to a newly introduced class of Barron-Bandlimited functions, for which it claims a dimension-independent Monte Carlo rate N^{-1/2} in the L2 norm and hence in the Fréchet metric.

Significance. The general reduction in Theorems 3.1 and 4.1 is a useful and clean idea: splitting the Fréchet sum into a finite leading part and a geometric tail, then using monotonicity of t/(1+t), converts any per-semi-norm rate into a sufficient width. The paper is also transparent about limitations: it proves that the bounded-growth condition (4.1) fails for the exponential spectral Barron space with Sobolev semi-norms (Propositions 4.2 and 4.3), and it does not rely on fitted parameters—all constants come from cited external rates. However, the headline bandlimited application is currently not established: Proposition 4.8 invokes a coefficient-budget theorem for a unit-coefficient approximation class, and the width formula of Theorem 4.9 depends on that unsupported rate. The manuscript also contains a self-referential appendix proposition. These issues are local and repairable, but they are load-bearing for the main example.

major comments (3)
  1. [Proposition 4.8 and Theorem 4.9] The approximation class bΣ_N defined in the proof of Proposition 4.8 consists of unit-coefficient sums Σ_{n=1}^N σ̂(⟨w_n, ·⟩+b_n), with no coefficients a_n. The proof invokes [57, Theorem 2] to assert an N^{-1/2} rate for this class. As stated in the cited theorem, however, the rate applies to networks with arbitrary coefficients a_n satisfying a bounded ℓ1 budget; the unit-coefficient class is a strict subset of that class for fixed N, and the infimum over a superset does not control the infimum over the subset. Moreover, unit coefficients force the ℓ1 budget to grow with N, so applying the cited theorem with budget N would not yield the claimed rate. Since Theorem 4.9 uses exactly r(N)=N^{-1/2} from Proposition 4.8, the dimension-independent Monte Carlo result is not established as written. The fix is local: add coefficients a_n with a constraint such as Σ|a_n| ≲ ||f||_{B*_1} to the definitions of Σ_N and bΣ_N, and then apply [57, Theorem 2] directly.
  2. [Theorem 3.1 and Theorem 4.1, proofs] In the proof of Theorem 3.1, after applying the approximation rate the manuscript writes inf dV ≤ (2-2^{-ℓ}) C_f M_ℓ r_ℓ(N)/(1+C_f M_ℓ r_ℓ(N)) + 2^{-ℓ} ≤ C_f M_ℓ r_ℓ(N) + 2^{-ℓ}. The last inequality is not valid in general: the factor (2-2^{-ℓ})/(1+C_f M_ℓ r_ℓ(N)) can exceed 1, for example when C_f M_ℓ r_ℓ(N) is small. With the threshold C_f M_ℓ r_ℓ(N) ≤ 2^{-ℓ} used in the proof, the correct bound is at most 3·2^{-ℓ}, so the displayed choice ℓ_ε = ⌈−log2 ε⌉+1 does not, as proven, guarantee an error below ε. The same issue appears in the proof of Theorem 4.1. This is repairable by taking a slightly larger ℓ_ε or a smaller threshold (e.g., replacing 2^{-ℓ} by 2^{-ℓ}/3), and the qualitative statements of the theorems survive, but the explicit width formulas need correction.
  3. [Appendix B, Proposition B.1] Proposition B.1 asserts that the weighted Gevrey class is either trivial or intractable to approximate, but its proof is self-referential: it states only that the proof 'could be found in the Appendix,' which is the same appendix containing the proposition. The surrounding discussion also contains an internal inconsistency: the differential-inequality argument |f'| ≤ M_1|f| implies that a C∞ function with a zero at a finite point vanishes on its connected component, which rules out compactly supported bump functions, whereas the preceding paragraph claims that in the non-analytic case β<1 the class may contain spatial concatenations of bump functions. The proposition should either be given a rigorous proof or removed, and the contradictory discussion should be reconciled.
minor comments (6)
  1. [Definition 4.7 and Proposition 4.8] Definition 4.7 defines B*_s(Ω) for s>1, but Proposition 4.8 states the result for f ∈ B*_1(Ω); if the intended class is B*_1, the definition and the theorem statements should be made consistent.
  2. [Theorem 4.9] In Theorem 4.9 the constant M_ℓ_ε is taken as ⟨Ω⟩^{ℓ_ε}, but for d-dimensional functions with supp f̂ ⊆ [−Ω,Ω]^d the correct Sobolev-growth constant is ⟨√d Ω⟩^{ℓ_ε} (up to a harmless constant); the dimension-free rate is unaffected, but the formula as written understates the constant.
  3. [Section 3.2, paragraph before Lemma 3.5] The claim that f̂ ∈ L1(e^{c|·|^β}) implies lim_{|ξ|→∞} |f̂(ξ)| e^{c|ξ|^β} = 0 and sup_ξ |f̂(ξ)| e^{c|ξ|^β} < ∞ is not valid for general L1 functions; the later inclusion Lemma 3.5 is proved by a different argument and does not depend on this claim, so the paragraph should be revised.
  4. [Section 1, after Theorem 3.1 discussion] The sentence claiming that this work is 'to the best of our knowledge the first such extension' to Fréchet spaces is contradicted by the cited works [12] and [33], which already treat neural networks in Fréchet spaces and Barron-type approximation in Fréchet metrics; the wording should be softened.
  5. [Lemma 3.4 and Definition 2.3] The notation S^{1,ℓ}(U) used in Lemma 3.4 is not defined; please use the notation S^p_{ω,ℓ}(U) from Definition 2.3 or state explicitly which p and weight are meant.
  6. [Proposition 4.3 proof] The displayed distributional Fourier transform of sin(nx) contains incorrect delta arguments (ξ/n instead of ξ∓n), and the subsequent change of variables is hard to follow; the computation should be rewritten with the stated Fourier convention.

Circularity Check

0 steps flagged · score 2.0 of 10

No circularity in the derivation chain: the main theorems are analytic reductions of externally supplied semi-norm rates to width bounds, and the Barron/Bandlimited applications rest on external cited rates; self-citations are background.

full rationale

The central derivation chain is not circular. Theorem 3.1 and Theorem 4.1 take as inputs a sequence of semi-norms, a growth or bounded-growth condition, and an assumed approximation rate r_l or r, and then compute a sufficient width N by inverting the rate. The Frechet metric is defined from the same semi-norms via (1.2), but the proofs do not simply restate that definition: they split the series, bound the tail by 2^{-l}, and use monotonicity of t/(1+t), so the metric error bound is a genuine consequence of the assumed rates rather than a definitional identity. The applications (Corollary 3.3, Corollary 3.6, Proposition 4.8, Theorem 4.9) rely on external approximation theorems [58, Theorem 2] and [57, Theorem 2] for the Barron-space rates; these are not self-citations, and no fitted parameter is renamed as a prediction. The authors' self-citations [2,3,4,5] appear as background for Gelfand-Shilov/Gevrey terminology and earlier variants of weighted spaces; none of the main proofs reduce to them. The manuscript itself flags the restrictive nature of the bounded-growth condition in Propositions 4.2 and 4.3, and Appendix B's Proposition B.1 contains a self-referential proof, but that is an omitted proof rather than circular reasoning. The sharpest concern is Proposition 4.8's invocation of [57, Theorem 2] for the unit-coefficient class bSigma_N; if that theorem requires a coefficient budget, the cited theorem may not apply, but this would be an unsupported external-citation or correctness gap, not a circular step. Overall, the paper is a reduction/meta-approximation framework over externally imported rates, with no significant circularity.

Assumptions & free parameters 2 free parameters · 5 assumptions · 2 invented entities

The central claim rests on prior approximation-rate theorems [57,58], on the standard Fréchet-metric construction, and on the newly introduced Barron-Bandlimited class. The metric weights 2^{-ℓ} and the growth constants are chosen by hand, and Proposition 4.2 assumes more about extensions than Lemma 3.4 proves.

free parameters (2)
  • Fréchet metric weights 2^{-ℓ} = 2^{-ℓ}
    Chosen by hand in Eq. (1.2). The sufficient width bounds N in Theorems 3.1 and 4.1 depend on these weights; any summable positive weight would yield a topologically equivalent metric but different quantitative bounds.
  • Growth constants M_ℓ = 1 for Sobolev norms on bounded U; ⟨Ω⟩^ℓ for bandlimited class
    The constants in conditions (3.1) and (4.1) are derived for the examples, not fitted to data. For Sobolev norms M_ℓ=1; for bandlimited functions M_ℓ=⟨Ω⟩^ℓ.
assumptions (5)
  • domain assumption Approximation rates from [58, Theorem 2] hold for exponential spectral Barron space with cosine networks
    Invoked in Corollary 3.3 and 3.6; the constants c_ℓ, C_ℓ are not tracked, so the width bound is existential.
  • domain assumption Approximation rate from [57, Theorem 2] holds for the frequency-domain activation \hat σ with σ in W^{s,1}
    Invoked in Proposition 4.8; the hypotheses of [57] on \hat σ (e.g., non-polynomial activation) are not verified in detail.
  • standard math The Fréchet metric (1.2) induces the same topology as the semi-norms
    Cited from [53, Remark 1.38(c)] and used throughout.
  • standard math Parseval's theorem and the support condition identify L2 error with frequency-domain L2 error
    Used in Proposition 4.8 to transfer the Monte Carlo rate from the Fourier side to the spatial side.
  • ad hoc to paper B_{β,c}(U) functions are real analytic for β ≥ 1, and all extensions are Barron functions over R^d
    Assumed in the proof of Proposition 4.2; Lemma 3.4 only establishes bounds for a minimizing sequence of extensions, not for all possible extensions. This is a gap in the proof.
invented entities (2)
  • Barron-Bandlimited space B^*_s(Ω)
    purpose: A target class that satisfies the bounded-growth condition (4.1) so that Theorem 4.1 applies with Sobolev semi-norms.
    Newly defined function class (Definition 4.7); it is a mathematical construction with no external falsifiable prediction.
  • Low-pass filtered inverse-Fourier network class Σ_N
    purpose: Approximation architecture that yields a dimension-free Monte Carlo rate for B^*_1(Ω).
    Defined in Proposition 4.8 and Theorem 4.9; it is not a standard feedforward network architecture, which limits its operator-learning relevance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Approximation Rates in Fr\'echet Metrics: Barron Spaces, Paley-Wiener Spaces, and Fourier Multipliers." pith.science (2026). https://pith.science/paper/IHH7BNAB

@misc{pith2026250104023,
  author       = {Pith},
  title        = {Pith review of: Approximation Rates in Fr\'echet Metrics: Barron Spaces, Paley-Wiener Spaces, and Fourier Multipliers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IHH7BNAB}},
  note         = {Machine review of arXiv:2501.04023}
}
read the original abstract

Operator learning is a recent development in the simulation of Partial Differential Equations (PDEs) by means of neural networks. The idea behind this approach is to learn the behavior of an operator, such that the resulting neural network is an (approximate) mapping in infinite-dimensional spaces that is capable of (approximately) simulating the solution operator governed by the PDE. In our work, we study some general approximation capabilities for linear differential operators by approximating the corresponding symbol in the Fourier domain. Analogous to the structure of the class of H\"ormander-Symbols, we consider the approximation with respect to a topology that is induced by a sequence of semi-norms. In that sense, we measure the approximation error in terms of a Fr\'echet metric, and our main result identifies sufficient conditions for achieving a predefined approximation error. Secondly, we then focus on a natural extension of our main theorem, in which we manage to reduce the assumptions on the sequence of semi-norms. Based on existing approximation results for the exponential spectral Barron space, we then present a concrete example of symbols that can be approximated well.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 38 canonical work pages

  1. [1]

    Uniform Approximation with Quadratic Neural Networks

    A. Abdeljawad, Uniform approximation with quadratic neural networks, Version Number: 3, 2022.doi: 10.48550/ARXIV.2201.03747

  2. [2]

    Pseudo-differentialcalculusinanisotropic gelfand–shilov setting,

    A.Abdeljawad,M.Cappiello,andJ.Toft,“Pseudo-differentialcalculusinanisotropic gelfand–shilov setting,”Integral Equations and Operator Theory, vol. 91, no. 3, p. 26, Jun. 2019.doi: 10.1007/s00020-019-2518-2

  3. [3]

    Liftings for ultra-modulation spaces, and one-parameter groups of gevrey-type pseudo-differential operators,

    A. Abdeljawad, S. Coriasco, and J. Toft, “Liftings for ultra-modulation spaces, and one-parameter groups of gevrey-type pseudo-differential operators,”Anal- ysis and Applications, vol. 18, no. 4, pp. 523–583, Jul. 2020.doi: 10.1142/ S0219530519500143

  4. [4]

    Abdeljawad and T

    A. Abdeljawad and T. Dittrich,Space-time approximation with shallow neural networks in fourier lebesgue spaces, Dec. 13, 2023. arXiv:2312.08461[cs]. 24

  5. [5]

    Abdeljawad and T

    A. Abdeljawad and T. Dittrich,Weighted sobolev approximation rates for neural networks on unbounded domains, Version Number: 1, 2024. doi: 10 . 48550 / ARXIV.2411.04108

  6. [6]

    Approximations with deep neural networks in sobolev time-space,

    A. Abdeljawad and P. Grohs, “Approximations with deep neural networks in sobolev time-space,”Analysis and Applications, vol. 20, no. 3, pp. 499–541, May

  7. [7]

    Integral representations of shallow neural net- work with rectified power unit activation function,

    A. Abdeljawad and P. Grohs, “Integral representations of shallow neural net- work with rectified power unit activation function,”Neural Networks, vol. 155, pp. 536–550, Nov. 2022.doi: 10.1016/j.neunet.2022.09.005

  8. [8]

    The cauchy problem for 3- evolution equations with data in gelfand–shilov spaces,

    A. Arias Junior, A. Ascanelli, and M. Cappiello, “The cauchy problem for 3- evolution equations with data in gelfand–shilov spaces,”Journal of Evolution Equations, vol. 22, no. 2, p. 33, Jun. 2022.doi: 10.1007/s00028-022-00764-z

Show all 66 references
  1. [9]

    Schrödinger-type equations in gelfand-shilov spaces,

    A. Ascanelli and M. Cappiello, “Schrödinger-type equations in gelfand-shilov spaces,” Journal de Mathématiques Pures et Appliquées, vol. 132, pp. 207–250, Dec. 2019. doi: 10.1016/j.matpur.2019.04.010

  2. [10]

    Universal approximation bounds for superpositions of a sig- moidal function,

    A. R. Barron, “Universal approximation bounds for superpositions of a sig- moidal function,” IEEE Transactions on Information Theory, vol. 39, no. 3, pp. 930–945, May 1993.doi: 10.1109/18.256500

  3. [11]

    Bateman and B

    H. Bateman and B. M. Project,Higher transcendental functions [Volumes I-III]. McGraw-Hill Book Company, 1953

  4. [12]

    Neuralnetworksinfréchetspaces,

    F.E.Benth,N.Detering,andL.Galimberti,“Neuralnetworksinfréchetspaces,” Annals of Mathematics and Artificial Intelligence, vol. 91, no. 1, pp. 75–103, Feb. 2023.doi: 10.1007/s10472-022-09824-z

  5. [13]

    Optimal approximation with sparsely connected deep neural networks,

    H. Bölcskei, P. Grohs, G. Kutyniok, and P. Petersen, “Optimal approximation with sparsely connected deep neural networks,”SIAM Journal on Mathematics of Data Science, vol. 1, no. 1, pp. 8–45, Jan. 2019.doi: 10.1137/18M118709X

  6. [14]

    Boullé and A

    N. Boullé and A. Townsend,A mathematical guide to operator learning, Version Number: 1, 2023.doi: 10.48550/ARXIV.2312.14688

  7. [15]

    Calvello, N

    E. Calvello, N. B. Kovachki, M. E. Levine, and A. M. Stuart,Continuum at- tention for neural operators, Version Number: 1, 2024.doi: 10.48550/ARXIV. 2406.06486

  8. [16]

    Propagation of exponential phase space sin- gularities for schrödinger equations with quadratic hamiltonians,

    E. Carypis and P. Wahlberg, “Propagation of exponential phase space sin- gularities for schrödinger equations with quadratic hamiltonians,”Journal of Fourier Analysis and Applications, vol. 23, no. 3, pp. 530–571, Jun. 2017.doi: 10.1007/s00041-016-9478-6

  9. [17]

    Universal approximation to nonlinear operators by neu- ralnetworkswitharbitraryactivationfunctionsanditsapplicationtodynamical systems,

    T. Chen and H. Chen, “Universal approximation to nonlinear operators by neu- ralnetworkswitharbitraryactivationfunctionsanditsapplicationtodynamical systems,” IEEE Transactions on Neural Networks, vol. 6, no. 4, pp. 911–917, Jul. 1995. doi: 10.1109/72.392253. 25

  10. [18]

    Characterizations of the gelfand-shilov spaces via fourier transforms,

    J. Chung, S.-Y. Chung, and D. Kim, “Characterizations of the gelfand-shilov spaces via fourier transforms,”Proceedings of the American Mathematical So- ciety, vol. 124, no. 7, pp. 2101–2108, 1996.doi: 10.1090/S0002- 9939- 96- 03291-1

  11. [19]

    Approximation by superpositions of a sigmoidal function,

    G. Cybenko, “Approximation by superpositions of a sigmoidal function,”Math- ematics of Control, Signals, and Systems, vol. 2, no. 4, pp. 303–314, Dec. 1989. doi: 10.1007/BF02551274

  12. [20]

    Generic bounds on the approximation error for physics-informed (and) operator learning,

    T. De Ryck and S. Mishra, “Generic bounds on the approximation error for physics-informed (and) operator learning,” Advances in Neural Information Processing Systems, vol. 35, pp. 10945–10958, 2022

  13. [21]

    The barron space and the flow-induced function spaces for neural network models,

    W. E, C. Ma, and L. Wu, “The barron space and the flow-induced function spaces for neural network models,”Constructive Approximation, vol. 55, no. 1, pp. 369–406, Feb. 2022.doi: 10.1007/s00365-021-09549-y

  14. [22]

    Kolmogorov width decay and poor approximators in machine learning: Shallow neural networks, random feature models and neural tangent kernels,

    W. E and S. Wojtowytsch, “Kolmogorov width decay and poor approximators in machine learning: Shallow neural networks, random feature models and neural tangent kernels,”Research in the Mathematical Sciences, vol. 8, no. 1, p. 5, Mar

  15. [23]

    I. M. Gel’fand and G. E. Šilov,Spaces of Fundamental and Generalized Func- tions (Generalized functions / I. M. Gel’fand, G. E. Shilov Volume 2), trans. by M. D. Friedman, A. Feinstein, and C. P. Peltzer. Providence, Rhode Island: AMS Chelsea Publishing, 2016, 261 pp

  16. [24]

    Solving high-dimensional partial differential equations using deep learning,

    J. Han, A. Jentzen, and W. E, “Solving high-dimensional partial differential equations using deep learning,” Proceedings of the National Academy of Sci- ences, vol. 115, no. 34, pp. 8505–8510, Aug. 21, 2018.doi: 10 . 1073 / pnas . 1718942115

  17. [25]

    L.Hörmander, Linear Partial Differential Operators.Berlin,Heidelberg:Springer Berlin Heidelberg, 1964.doi: 10.1007/978-3-662-30724-3

  18. [26]

    L.Hörmander, The Analysis of Linear Partial Differential Operators I(Grundlehren der mathematischen Wissenschaften), red. by M. Artinet al.Berlin, Heidelberg: Springer Berlin Heidelberg, 1998, vol. 256.doi: 10.1007/978-3-642-96750-4

  19. [27]

    Berlin, Heidelberg: Springer Berlin Heidelberg, 2007.doi: 10.1007/978-3-540-49938-1

    L.Hörmander, The Analysis of Linear Partial Differential Operators III: Pseudo- Differential Operators(Classics in Mathematics). Berlin, Heidelberg: Springer Berlin Heidelberg, 2007.doi: 10.1007/978-3-540-49938-1

  20. [28]

    Approximation capabilities of multilayer feedforward networks,

    K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural networks, vol. 4, no. 2, pp. 251–257, 1991, Publisher: Elsevier

  21. [29]

    Basis operator network: A neural network-based model for learning nonlinear operators via neural basis,

    N. Hua and W. Lu, “Basis operator network: A neural network-based model for learning nonlinear operators via neural basis,”Neural Networks, vol. 164, pp. 21–37, Jul. 2023.doi: 10.1016/j.neunet.2023.04.017. 26

  22. [30]

    D. Z. Huang, N. H. Nelsen, and M. Trautner,An operator learning perspective on parameter-to-observable maps, Version Number: 2, 2024.doi: 10 .48550 / ARXIV.2402.06031

  23. [31]

    MIONet: Learning multiple-input operators via tensor product,

    P. Jin, S. Meng, and L. Lu, “MIONet: Learning multiple-input operators via tensor product,”SIAM Journal on Scientific Computing, vol. 44, no. 6, A3490– A3514, Dec. 2022.doi: 10.1137/22M1477751

  24. [32]

    A characterization of real analytic functions,

    H. Komatsu, “A characterization of real analytic functions,”Proceedings of the Japan Academy, Series A, Mathematical Sciences, vol. 36, no. 3, Jan. 1, 1960. doi: 10.3792/pja/1195524081

  25. [33]

    Two-layer neural networks with values in a banach space,

    Y. Korolev, “Two-layer neural networks with values in a banach space,”SIAM Journal on Mathematical Analysis, vol. 54, no. 6, pp. 6358–6389, Dec. 2022. doi: 10.1137/21M1458144

  26. [34]

    Neural operator: Learning maps between function spaces with applications to PDEs,

    N. Kovachki et al., “Neural operator: Learning maps between function spaces with applications to PDEs,” Journal of Machine Learning Research, vol. 24, no. 89, pp. 1–97, 2023

  27. [35]

    N. B. Kovachki, S. Lanthaler, and A. M. Stuart,Operator learning: Algorithms and analysis, Version Number: 1, 2024.doi: 10.48550/ARXIV.2402.15715

  28. [36]

    Operator learning with PCA-net: Upper and lower complexity bounds,

    S. Lanthaler, “Operator learning with PCA-net: Upper and lower complexity bounds,” Journal of Machine Learning Research, vol. 24, no. 318, pp. 1–67, 2023

  29. [37]

    Lanthaler, Z

    S. Lanthaler, Z. Li, and A. M. Stuart, Nonlocality and nonlinearity implies universality in operator learning, Version Number: 2, 2023. doi: 10 . 48550 / ARXIV.2304.13221

  30. [38]

    ErrorestimatesforDeepONets: A deep learning framework in infinite dimensions,

    S.Lanthaler,S.Mishra,andG.E.Karniadakis,“ErrorestimatesforDeepONets: A deep learning framework in infinite dimensions,”Transactions of Mathematics and Its Applications, vol. 6, no. 1, tnac001, Mar. 8, 2022.doi: 10.1093/imatrm/ tnac001

  31. [39]

    Multilayer feedforward net- works with a nonpolynomial activation function can approximate any function,

    M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken, “Multilayer feedforward net- works with a nonpolynomial activation function can approximate any function,” Neural Networks, vol. 6, no. 6, pp. 861–867, Jan. 1993.doi: 10.1016/S0893- 6080(05)80131-5

  32. [40]

    Two-layer networks with the ReLU$^k$ activation function: Barron spaces and derivative approximation,

    Y. Li, S. Lu, P. Mathé, and S. V. Pereverzev, “Two-layer networks with the ReLU$^k$ activation function: Barron spaces and derivative approximation,” Numerische Mathematik, Nov. 23, 2023.doi: 10.1007/s00211-023-01384-6

  33. [41]

    Fourier neural operator for parametric partial differential equa- tions,

    Z. Li et al., “Fourier neural operator for parametric partial differential equa- tions,” 2020. arXiv:2010.08895. 27

  34. [42]

    A priori generalization error analysis of two-layer neural net- works for solving high dimensional schrödinger eigenvalue problems,

    J. Lu and Y. Lu, “A priori generalization error analysis of two-layer neural net- works for solving high dimensional schrödinger eigenvalue problems,”Commu- nications of the American Mathematical Society, vol. 2, no. 1, pp. 1–21, Jan. 31,

  35. [43]

    Deepnetworkapproximationforsmooth functions,

    J.Lu,Z.Shen,H.Yang,andS.Zhang,“Deepnetworkapproximationforsmooth functions,” SIAM Journal on Mathematical Analysis, vol. 53, no. 5, pp. 5465– 5506, Jan. 2021.doi: 10.1137/20M134695X

  36. [44]

    Learning nonlinear operators via DeepONet based on the universal approximation theorem of op- erators,

    L. Lu, P. Jin, G. Pang, Z. Zhang, and G. E. Karniadakis, “Learning nonlinear operators via DeepONet based on the universal approximation theorem of op- erators,” Nature Machine Intelligence, vol. 3, no. 3, pp. 218–229, Mar. 18, 2021. doi: 10.1038/s42256-021-00302-5

  37. [45]

    doi: https://doi.org/10.1090/cams/5

  38. [46]

    Uniform approximation rates and metric en- tropyofshallowneuralnetworks,

    L. Ma, J. W. Siegel, and J. Xu, “Uniform approximation rates and metric en- tropyofshallowneuralnetworks,” Research in the Mathematical Sciences,vol.9, no. 3, p. 46, Sep. 2022.doi: 10.1007/s40687-022-00346-y

  39. [47]

    Uniform approximation by neural networks,

    Y. Makovoz, “Uniform approximation by neural networks,”Journal of Approx- imation Theory, vol. 95, no. 2, pp. 215–228, Nov. 1998.doi: 10.1006/jath. 1997.3217

  40. [48]

    Two-layer neural networks for partial differential equa- tions: Optimization and generalization theory,

    T. Luo and H. Yang, “Two-layer neural networks for partial differential equa- tions: Optimization and generalization theory,” 2020. arXiv:2006.15733

  41. [49]

    Approximation by superposition of sigmoidal and radial basis functions,

    H. Mhaskar and C. A. Micchelli, “Approximation by superposition of sigmoidal and radial basis functions,”Advances in Applied Mathematics, vol. 13, no. 3, pp. 350–373, Sep. 1992.doi: 10.1016/0196-8858(92)90016-P

  42. [50]

    Deep ReLU networks overcome the curse of dimensionality for generalized bandlimited functions,

    H. Montanelli, H. Yang, and Q. Du, “Deep ReLU networks overcome the curse of dimensionality for generalized bandlimited functions,”Journal of Computa- tional Mathematics, vol. 39, no. 6, pp. 801–815, Jan. 1, 2021.doi: 10.4208/ jcm.2007-m2019-0239

  43. [51]

    Neural networks for functional approximation and system identification,

    H. N. Mhaskar and N. Hahm, “Neural networks for functional approximation and system identification,” Neural Computation, vol. 9, no. 1, pp. 143–159, Jan. 1, 1997.doi: 10.1162/neco.1997.9.1.143

  44. [52]

    Rudin, Real and complex analysis, 3rd ed

    W. Rudin, Real and complex analysis, 3rd ed. New York: McGraw-Hill, 1987, 416 pp

  45. [53]

    Rudin, Functional analysis(International series in pure and applied math- ematics), 2nd ed

    W. Rudin, Functional analysis(International series in pure and applied math- ematics), 2nd ed. New York: McGraw-Hill, 1991, 424 pp

  46. [54]

    Prasthofer, T

    M. Prasthofer, T. De Ryck, and S. Mishra,Variable-input deep operator net- works, Version Number: 1, 2022.doi: 10.48550/ARXIV.2205.11404

  47. [55]

    Deep operator network approximation rates for lipschitz operators,

    C. Schwab, A. Stein, and J. Zech, “Deep operator network approximation rates for lipschitz operators,”arXiv preprint, 2023

  48. [56]

    Optimal approximation rates for deep ReLU neural networks on sobolev and besov spaces,

    J. W. Siegel, “Optimal approximation rates for deep ReLU neural networks on sobolev and besov spaces,”Journal of Machine Learning Research, vol. 24, no. 357, pp. 1–52, 2023

  49. [57]

    Howdo infinitewidthbounded norm networks look in function space?

    P.Savarese,I. Evron, D. Soudry,and N. Srebro,“Howdo infinitewidthbounded norm networks look in function space?” In Proceedings of the Thirty-Second Conference on Learning Theory, PMLR, Jun. 25, 2019, pp. 2667–2690. 28

  50. [58]

    High-order approximation rates for shallow neural net- works with cosine and ReLU activation functions,

    J. W. Siegel and J. Xu, “High-order approximation rates for shallow neural net- works with cosine and ReLU activation functions,”Applied and Computational Harmonic Analysis, vol. 58, pp. 1–26, May 2022.doi: 10.1016/j.acha.2021. 12.005

  51. [59]

    Sharp bounds on the approximation rates, metric en- tropy, and n-widths of shallow neural networks,

    J. W. Siegel and J. Xu, “Sharp bounds on the approximation rates, metric en- tropy, and n-widths of shallow neural networks,”Foundations of Computational Mathematics, Nov. 9, 2022.doi: 10.1007/s10208-022-09595-3

  52. [60]

    Approximation rates for neural networks with general activation functions,

    J. W. Siegel and J. Xu, “Approximation rates for neural networks with general activation functions,”Neural Networks, vol. 128, pp. 313–321, Aug. 2020.doi: 10.1016/j.neunet.2020.05.019

  53. [61]

    Functional analytic characterizations of the gelfand-shilov spaces $s_\alpha^\beta$,

    S. Van Eijndhoven, “Functional analytic characterizations of the gelfand-shilov spaces $s_\alpha^\beta$,” Indagationes Mathematicae (Proceedings), vol. 90, no. 2, pp. 133–144, Jun. 1987.doi: 10.1016/S1385-7258(87)80035-5

  54. [62]

    Near-optimal deep neural network approximation for ko- robov functions with respect to l p and h 1 norms,

    Y. Yang and Y. Lu, “Near-optimal deep neural network approximation for ko- robov functions with respect to l p and h 1 norms,”Neural Networks, vol. 180, p. 106702, Dec. 2024.doi: 10.1016/j.neunet.2024.106702

  55. [63]

    Subedi and A

    U. Subedi and A. Tewari, Error bounds for learning fourier linear operators, Aug. 16, 2024. arXiv:2408.09004[cs,math,stat]

  56. [66]

    Optimal approximation of continuous functions by very deep ReLU networks,

    D. Yarotsky, “Optimal approximation of continuous functions by very deep ReLU networks,” inConference on learning theory, PMLR, 2018, pp. 639–649. Appendix A. Observations on the Exponential Spectral Barron Space Proposition Appendix A.1. The L2(U )-unit ball is unbounded inBβ...

  57. [2021]

    doi: 10.1007/s40687-020-00233-4

  58. [2022]

    doi: 10.1142/S0219530522500014

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.