REVIEW 2 major objections 4 minor 2 cited by
This paper argues that the sample complexity of learning an operator from point evaluations is governed by the operator's regularity, proving that holomorphic operators admit algebraic error rates whereas Lipschitz or C^k operators cannot.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A survey of operator learning theory showing holomorphy gives fast sample-complexity rates, general smoothness gives a polylogarithmic barrier, and FNO-approximable classes cap out at n^{-1/2}.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection A transparent survey of operator learning sample-complexity bounds; the FNO exponent ceiling in Thm. 5 rests on an unverified lemma transfer. the 2 major comments →
A short tour of operator learning theory: Convergence rates, statistical limits, and open questions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the paper's own terms, the central discovery is a hierarchy of minimax learnability for operator classes. For the unit ball in the Lipschitz or C^k spaces, the best achievable worst-case error over all methods using n point evaluations decays only polylogarithmically, so no algebraic rate is possible (Thm. 3). For holomorphic operators with l^p-summable parameter sequences, the minimax error is n^{-(1/p-1/2)} up to log factors, a rate that can be made arbitrarily close to n^{-1/2} as p→0 and is attained by an empirical risk minimizer based on compressed sensing (Thms. 2 and 4). For operators that are efficiently approximated by Fourier Neural Operators at rate α, the optimal algebraic min
What carries the argument
The central objects are the nonlinear sampling n-width s_n(K)_X, which records the best worst-case error achievable with n point evaluations over a class K, and the regularity classes themselves: the holomorphic classes H(b) built from Bernstein polyellipses with l^p-summability, and the FNO-approximable classes K_alpha^FNO. The upper bounds are obtained by balancing approximation error from finite-dimensional encoding/decoding with statistical error from finite samples, using empirical process theory (for ReLU networks) and compressed sensing (for handcrafted tanh networks). The lower bounds combine n-width arguments with concrete constructions of hard operator families. The key mechanism i
Load-bearing premise
The load-bearing premise is that the empirical risk minimization problem (3) has a solution and that training can find it: the theorems are existential for global minimizers, and no algorithm is proven to reach them, so the sample-complexity rates apply only to an idealized optimization.
What would settle it
Find a concrete learning method—neural or otherwise—that achieves an algebraic error rate for the unit Lipschitz ball from n point evaluations, contradicting Theorem 3; or construct a holomorphic class with l^p-summable parameters whose minimax rate is strictly slower than n^{-(1/p-1/2)}, contradicting Theorem 4.
If this is right
- For holomorphic operators with l^p-summable parameter sequences and negligible encoder/decoder error, empirical risk minimization achieves minimax-optimal error rates n^{-(1/p - 1/2)} in the absence of noise, which are algebraic and can be made arbitrarily fast as p decreases (Thms. 2 and 4).
- No learning method can achieve algebraic sample complexity for the unit Lipschitz or C^k ball from point evaluations; the best possible worst-case error decays only polylogarithmically (Thm. 3).
- For operators that are well approximated by Fourier Neural Operators at rate α, the optimal algebraic minimax rate exponent is between 1/(2+16/α) and 1/2, so algebraic rates are possible but capped at n^{-1/2} (Thm. 5).
- The minimax errors for the holomorphic and FNO classes are both bounded above by the minimax error for the Lipschitz class, giving a clean separation between regularity classes (paper's §3 discussion).
Where Pith is reading between the lines
- A practical consequence of the separation is that estimating the analytic regularity of a solution map (e.g., from parameter-to-solution maps of PDEs) before choosing an architecture would be more valuable than architecture search; holomorphy justifies expecting fast convergence.
- The lower bounds assume point-evaluation measurements; if linear functionals (e.g., Fourier coefficients) are allowed as measurements, the Lipschitz lower bound may not apply, and the sample complexity could change—this is not explored in the survey.
- The theorem results are existential for global minimizers of a nonconvex problem; without proofs that gradient-based training can reach these minimizers, the theoretical rates may not be realized in practice, an open problem the paper acknowledges.
- One could test the FNO upper bound by training FNOs on holomorphic operators: if the observed error exponent exceeds 1/2 for large α, the true minimax exponent might be higher than the survey's guess.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey paper presents a unified, notationally consistent overview of recent results in operator learning theory with emphasis on sample complexity. Section 2 reviews two empirical-risk-minimization bounds for holomorphic operators: Theorem 1, adapted from Reinhardt–Wang–Zech, gives an expectation error bound with a rate approaching n^{-1/2} for deep ReLU networks; Theorem 2, adapted from Adcock–Dexter–Moraga, gives a faster-than-Monte-Carlo algebraic rate for a handcrafted tanh-network class under polynomial truncation-error decay and bounded noise. Section 3 surveys minimax results: Theorem 3 gives a polylogarithmic lower bound for C^k and Lipschitz operator classes from point evaluations; Theorem 4 gives lower bounds for holomorphic classes showing rates of order n^{-(1/p-1/2)}; Theorem 5 gives two-sided bounds for FNO-approximable classes, with optimal algebraic exponent between roughly 1/2 and 1/2; and Theorem 6 extends the story to a noisy sampling setting. The paper closes with several open problems. The authors are transparent that they are re-presenting results from cited papers in a unified notation, with a partial loss of generality.
Significance. If the results are correct as presented, the survey provides a valuable synthesis: holomorphy yields arbitrarily fast algebraic minimax rates, Lipschitz/C^k regularity yields only polylogarithmic rates, and FNO-approximability yields algebraic but no-better-than-n^{-1/2} rates. This gives a clear organizing principle, essentially 'regularity governs sample complexity.' The explicit attribution of each theorem to a cited source is a strength, as is the careful separation of upper and lower bounds and the candid discussion of open problems. The paper does not prove new theorems, but a survey does not need to; its contribution is the coherent framing. The weakest point is a single load-bearing verification in Theorem 5, discussed below; if that is repaired, the survey would be a useful reference for the operator-learning community.
major comments (2)
- [§3, Theorem 5 and the proof paragraph after Eq. (10)] The upper bound β*≤1/2 is load-bearing for the FNO half of the paper's central conclusion, yet the proof rests on the sentence that 'perusal' of the proof of [16, Lem. 3.22] shows that [16, Eqn. (3.30)] remains valid with K^α_FNO(K) in place of U^{α,∞}_{ℓ,NO}. This is not a formality: the class K^α_FNO(K) defined in Eq. (9) involves a specific FNO architecture, parameter norm bound exp(m), and sup_{m} m^α e_K(f,N_m^{FNO}) approximation error, while the class in [16] may measure approximation differently. The paper gives no argument that the metric-entropy lower bound or the relevant construction transfers. Please supply a proof or a precise reduction—for example, show that U^{α,∞}_{ℓ,NO} embeds into K^α_FNO(K) with explicit parameter choices, or state and prove a lemma establishing the lower bound directly for K^α_FNO(K). Without this, Eq. (10)'s upper bound is unsupported and the 'no fa
- [§3, Theorem 4 and the paragraph following it] The displayed theorem contains only lower bounds on s_n(H(b)∘ι), but the text states that the rate is 'optimal, up to log factors.' This optimality conclusion is not a consequence of the theorem as stated; it requires a matching upper bound, which is only indirectly available through Theorem 2 under extra hypotheses (e.g., Eq. (6), σ=0). Please state the matching upper bound explicitly, or rephrase the optimality claim to identify precisely which theorem supplies the matching upper bound and under which assumptions. This would prevent a reader from attributing to Theorem 4 an assertion it does not contain.
minor comments (4)
- [§2, Theorems 1–2] Both theorems are stated for global minimizers of the nonconvex ERM objective (3). Existence is asserted or guaranteed with high probability, but no algorithm is claimed to find such minimizers. The paper acknowledges this only in the §4 open-problem discussion. Since the abstract promises convergence rates for empirical risk minimization, add a remark in §2 clarifying the oracle nature of these results and referring the reader to the open problems.
- [§3, Eq. (9)] The notation e_K(f,𝒩) in the definition of Γ^α_FNO is ambiguous because 𝒩 is not defined at that point; it should be e_K(f,𝒩_m^{FNO}). Also, the rendering of the map ι as U→R^N should be U→ℝ^ℕ, and the inclusion H_{0,1}=[−1,1]^ℕ⊆R(b) is meant in ℓ^∞, not ℓ^2; clarify the ambient space.
- [§3, Theorem 6, Eq. (12)] The displayed chain contains three inequalities, but the following text refers only to 'the first' and 'the second' of them. Label the inequalities (e.g., (12a), (12b), (12c)) to make the explanation unambiguous.
- [§2, Theorem 1, Eq. (4)] The additive τ in the right-hand side means that for fixed τ>0 the displayed bound does not tend to zero as n→∞. The text's phrase 'approximate Monte Carlo rate' is intended to address this, but the dependence of c on τ is not stated. Please write c_τ or add a sentence clarifying that c depends on τ and other fixed parameters, and that the τ term is to be sent to zero in the limiting-rate statement.
Circularity Check
No significant circularity: the load-bearing theorems are imported from external groups and are not equivalent to the survey's inputs.
full rationale
The paper is a survey. Its main theorems are quoted from external groups: Thm. 1 from Reinhardt-Wang-Zech [35], Thm. 2 from Adcock-Dexter-Moraga [4], Thms. 3 and 5 from Kovachki-Lanthaler-Mhaskar [20] and Grohs-Lanthaler-Trautner [16], and Thm. 6 from Adcock-Maier-Parhi [6]. The survey adds no derivation in which an output quantity is defined in terms of, or fitted to, the quantity it later predicts. The ERM upper bounds and minimax lower bounds are independent statements; matching them in Thm. 4 and the Discussion is a synthesis, not a construction. The class K^alpha_FNO is defined by FNO approximation rates, not by sample complexity, so the bound beta* <= 1/2 is a genuine transfer from metric-entropy estimates in [16], not a tautology. The authors' self-citations ([13], [15], [24], [30], [31]) appear in background, context, and one illustrative example; none supplies a central conclusion. In particular, [13] is cited only as an explicit measure example satisfying hypotheses already provided by [16, Prop. 3.5], so it is not load-bearing. The only load-bearing concern is the unshown transfer in Theorem 5: the paper says 'perusal of the proof of [16, Lem. 3.22] show that [16, Eqn. (3.30)] remains valid with K^alpha_FNO in place of U^{alpha,infty}_{ell,NO}', which is a missing verification at a joint, not a circular reduction. It should be weighed as correctness risk, not circularity. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' own prior work, and no ansatz is smuggled in via self-citation. The derivation chain is not equivalent to its inputs, so the circularity burden is low.
Axiom & Free-Parameter Ledger
free parameters (2)
- tau in Theorem 1 (slack exponent) =
arbitrarily small > 0 (not estimated from data)
- delta in Theorem 2 (width slack) =
arbitrarily small > 0 (not estimated from data)
axioms (7)
- domain assumption Assumption 1: existence of boundedly invertible linear maps E_infty: U -> l2 and D_infty: l2 -> V, with truncated encoder-decoder architecture D_q ∘ g ∘ E_d.
- domain assumption Holomorphy: G admits a holomorphic extension over O containing E_infty^{-1}(H_{r,R}) (Thm. 1), or G∘E_infty^{-1} ∈ H(b) with b ∈ l^p_M on Bernstein polyellipses (Thm. 2).
- domain assumption Spectral decay: supp(ρ) ⊆ E_infty^{-1}(H_{r,R}) and sup norm condition in Thm. 1; truncation errors decay as s^{-γ} and s^{-ν+1/2} in (6a)–(6b); covariance eigenvalues satisfy inf_j j^{2ϑ} λ_j > 0 or λ_j ≍ j^{-2ϑ} in Thms. 3 and 6.
- domain assumption Measure conditions: (E_infty)_#ρ is quasi-uniform on H_{0,1} and (R_s∘E_infty)_#ρ is absolutely continuous w.r.t. uniform measure (Thm. 2); ρ is Gaussian in Thms. 3 and 6; ρ is supported on a compact convex infinite-dimensional K in Thm. 5.
- domain assumption Noise models: i.i.d. subgaussian noise (Thm. 1), bounded noise (Thm. 2), Gaussian noise (Thm. 6).
- ad hoc to paper Existence and attainability of a global minimizer of the ERM objective (3).
- domain assumption Minimax information model: the learner is restricted to n point evaluations of the target operator with arbitrary decoding, as encoded in Map_n and the nonlinear sampling n-width (8).
Cite this review
Pith. "Pith review of A short tour of operator learning theory: Convergence rates, statistical limits, and open questions." pith.science (2026). https://pith.science/paper/PY5VZ5XN
@misc{pith2026260300819,
author = {Pith},
title = {Pith review of: A short tour of operator learning theory: Convergence rates, statistical limits, and open questions},
year = {2026},
howpublished = {\url{https://pith.science/paper/PY5VZ5XN}},
note = {Machine review of arXiv:2603.00819}
}
read the original abstract
This paper surveys recent developments at the intersection of operator learning, statistical learning theory, and approximation theory. First, it reviews error bounds for empirical risk minimization with a focus on holomorphic operators and neural network approximations. Next, it illustrates fundamental performance limits in terms of sample size by adopting a minimax perspective and considering various notions of regularity beyond holomorphy. The paper ends with a discussion on the interplay between these two perspectives and related open questions.
Forward citations
Cited by 2 Pith papers
-
One Operator for Many Densities: Amortized Approximation of Conditioning by Neural Operators
A single neural operator can approximate the map from arbitrary joint densities to their conditionals, backed by new continuity results and illustrated on Gaussian mixtures.
-
One Operator for Many Densities: Amortized Approximation of Conditioning by Neural Operators
A single neural operator can approximate the map from joint densities to conditional densities to arbitrary accuracy, with a proof based on continuity of the conditioning operator and a demonstration on Gaussian mixtures.
Reference graph
Works this paper leans on
-
[1]
Adcock, S
B. Adcock, S. Brugiapaglia, N. Dexter, and S. Moraga. Learning smooth functions in high dimensions: From sparse polynomials to deep neural networks. In S. Mishra and A. Townsend, editors,Handb. Numer. Anal., volume 25, pages 1–52. Elsevier, 2024
2024
-
[2]
Adcock, S
B. Adcock, S. Brugiapaglia, N. Dexter, and S. Moraga. Near-optimal learning of Banach-valued, high-dimensional functions via deep neural networks.Neural Netw., 181, 2025
2025
-
[3]
Adcock, S
B. Adcock, S. Brugiapaglia, and C. G. Webster.Sparse Polynomial Approximation of High- Dimensional Functions, volume 25. SIAM, 2022
2022
-
[4]
Adcock, N
B. Adcock, N. Dexter, and S. Moraga. Optimal deep learning of holomorphic operators between Banach spaces. InAdv. Neural Inf. Process. Syst., volume 37, pages 27725–27789, 2024
2024
-
[5]
B. Adcock, M. Griebel, and G. Maier. The sample complexity of learning Lipschitz operators with respect to Gaussian measures.preprint arXiv:2410.23440, 2024
Pith/arXiv arXiv 2024
- [6]
-
[7]
Bhattacharya, B
K. Bhattacharya, B. Hosseini, N. B. Kovachki, and A. M. Stuart. Model reduction and neural networks for parametric PDEs.The SMAI J. Comput. Math., 7:121–157, 2021
2021
-
[8]
Boull ´e and A
N. Boull ´e and A. Townsend. A mathematical guide to operator learning. In S. Mishra and A. Townsend, editors,Handb. Numer. Anal., volume 25, pages 83–125. Elsevier, 2024
2024
-
[9]
Chen and H
T. Chen and H. Chen. Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems.IEEE Trans. Neural Netw., 6(4):911–917, 1995
1995
-
[10]
Cohen and R
A. Cohen and R. DeVore. Approximation of high-dimensional parametric PDEs.Acta Numer., 24:1–159, 2015
2015
-
[11]
Cohen, C
A. Cohen, C. Schwab, and J. Zech. Shape holomorphy of the stationary Navier–Stokes equations. SIAM J. Math. Anal., 50(2):1720–1752, 2018
2018
-
[12]
G. Cybenko. Approximation by superpositions of a sigmoidal function.Math. Control. Signals Syst., 2(4):303–314, 1989
1989
-
[13]
M. V. de Hoop, N. B. Kovachki, M. Lassas, and N. H. Nelsen. Extension and neural operator approximation of the electrical impedance tomography inverse map.preprint arXiv:2511.20361, 2025
arXiv 2025
-
[14]
R. DeVore, R. D. Nowak, R. Parhi, G. Petrova, and J. W. Siegel. Optimal recovery meets minimax estimation.preprint arXiv:2502.17671, 2025. 12 Simone Brugiapaglia, Nicola Rares Franco, Nicholas H. Nelsen
Pith/arXiv arXiv 2025
-
[15]
N. R. Franco and S. Brugiapaglia. A practical existence theorem for reduced order models based on convolutional autoencoders.Found. Data Sci., 7(1):72–98, 2025
2025
-
[16]
P. Grohs, S. Lanthaler, and M. Trautner. Theory-to-practice gap for neural networks and neural operators.preprint arXiv:2503.18219, 2025
Pith/arXiv arXiv 2025
-
[17]
G¨ uhring and M
I. G¨ uhring and M. Raslan. Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Netw., 134:107–130, 2021
2021
-
[18]
Gy ¨orfi, M
L. Gy ¨orfi, M. Kohler, A. Krzy˙ zak, and H. Walk.A distribution-free theory of nonparametric regression. Springer New York, 2002
2002
-
[19]
K. Hornik. Approximation capabilities of multilayer feedforward networks.Neural Netw., 4(2):251–257, 1991
1991
-
[20]
N. B. Kovachki, S. Lanthaler, and H. Mhaskar. Data complexity estimates for operator learning. preprint arXiv:2405.15992, 2024
Pith/arXiv arXiv 2024
-
[21]
N. B. Kovachki, S. Lanthaler, and A. M. Stuart. Operator learning: Algorithms and analysis. In S. Mishra and A. Townsend, editors,Handb. Numer. Anal., volume 25, pages 419–467. Elsevier, 2024
2024
-
[22]
D. Krieg and M. Ullrich. Approximation of functions: Optimal sampling and complexity.preprint arXiv:2602.02066, 2026
Pith/arXiv arXiv 2026
-
[23]
Lanthaler, S
S. Lanthaler, S. Mishra, and G. E. Karniadakis. Error estimates for DeepONets: A deep learning framework in infinite dimensions.Trans. Math. Appl., 6(1), 2022
2022
-
[24]
Lanthaler and N
S. Lanthaler and N. H. Nelsen. Error bounds for learning with vector-valued random features. In Adv. Neural Inf. Process. Syst., volume 36, pages 71834–71861, 2023
2023
-
[25]
Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. M. Stuart, and A. Anand- kumar. Fourier neural operator for parametric partial differential equations.Int. Conf. Learn. Represent., 2021
2021
-
[26]
H. Liu, H. Yang, M. Chen, T. Zhao, and W. Liao. Deep nonparametric estimation of operators between infinite dimensional spaces.J. Mach. Learn. Res., 25(24):1–67, 2024
2024
-
[27]
L. Lu, P. Jin, G. Pang, Z. Zhang, and G. E. Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators.Nature Mach. Intell., 3(3):218–229, 2021
2021
-
[28]
Marcati, J
C. Marcati, J. A. Opschoor, P. C. Petersen, and C. Schwab. Exponential ReLU neural network approximation rates for point and edge singularities.Found. Comput. Math., 23(3):1043–1127, 2023
2023
-
[29]
H. N. Mhaskar. Neural networks for optimal approximation of smooth and analytic functions. Neural Comput., 8(1):164–177, 1996
1996
-
[30]
N. H. Nelsen and A. M. Stuart. Operator learning using random features: A tool for scientific computing.SIAM Rev., 66(3):535–571, 2024
2024
-
[31]
N. H. Nelsen and Y. Yang. Operator learning meets inverse problems: A probabilistic perspective. preprint arXiv:2508.20207, 2025
arXiv 2025
-
[32]
Novak and H
E. Novak and H. Wo´ zniakowski.Tractability of Multivariate Problems. Volume I: Linear Infor- mation, volume 6 ofEMS Tracts Math.EMS, Z¨ urich, Switzerland, 2008
2008
-
[33]
Parhi and B
R. Parhi and B. Adcock. Upper bounds on averaged sampling numbers for general model classes. In2025 Int. Conf. Sampl. Theory Appl., pages 1–5. IEEE, 2025
2025
-
[34]
Petersen and F
P. Petersen and F. Voigtlaender. Optimal approximation of piecewise smooth functions using deep ReLU neural networks.Neural Netw., 108:296–330, 2018
2018
-
[35]
N. Reinhardt, S. Wang, and J. Zech. Statistical learning theory for neural operators.preprint arXiv:2412.17582, 2024
Pith/arXiv arXiv 2024
-
[36]
Schwab and J
C. Schwab and J. Zech. Deep learning in high dimension: Neural network expression rates for analytic functions in𝑙 2(R𝑑,𝛾 𝑑).SIAM/ASA J. Uncertain. Quantif., 11(1):199–234, 2023
2023
-
[37]
J. W. Siegel. Optimal approximation rates for deep ReLU neural networks on Sobolev and Besov spaces.J. Mach. Learn. Res., 24(357):1–52, 2023
2023
-
[38]
J. W. Siegel and J. Xu. Sharp bounds on the approximation rates, metric entropy, and𝑛-widths of shallow neural networks.Found. Comput. Math., 24(2):481–537, 2024
2024
-
[39]
A. v. d. Vaart and J. A. Wellner. Empirical processes. InWeak Convergence and Empirical Processes: With Applications to Statistics, pages 127–384. Springer, 2023
2023
-
[40]
Yarotsky
D. Yarotsky. Error bounds for approximations with deep ReLU networks.Neural Netw., 94:103– 114, 2017
2017
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.