{"id":"58708a06-5899-47c2-8827-c7feca1fcfc8","arxiv_id":"2603.00819","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A survey of operator learning theory showing holomorphy gives fast sample-complexity rates, general smoothness gives a polylogarithmic barrier, and FNO-approximable classes cap out at n^{-1/2}.","lead":"This survey organizes recent mathematical results on how many training examples are needed to learn operators—maps between infinite-dimensional spaces—with neural networks. It shows that very smooth (holomorphic) operators can be learned at fast algebraic rates, while merely differentiable operators suffer a near-logarithmic sample-complexity barrier.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Thm. 5's upper bound on the FNO-approximable exponent rests on an unshown transfer of [16, Lem. 3.22] to the class K^α_FNO; this is the least-secure load-bearing step.","rationale":"The reader's weakest assumption concerns nonconvex ERM and the gap between global minimizers and trainable optima. That is a real practical caveat, and the survey does acknowledge it, especially in the discussion of handcrafted architectures and in the §4 open problems. However, it does not directly threaten the minimax claims, which are about information-theoretic limits. The more load-bearing spot is Thm. 5: the paper's quantitative upper bound of 1/2 on the FNO-approximable exponent is a key pillar of the 'architecture alone is not enough' narrative, and the only provided support is a terse 'perusal' claim that a lemma from another preprint transfers to the paper's class K^α_FNO. If that transfer fails, the central comparison between FNO-approximable classes and smooth classes is left without its upper ceiling. The paper may well be correct, but the missing verification should be supplied or the claim should be explicitly flagged as relying on an unverified transfer. Hence I recommend CONDITIONAL acceptance rather than unconditional acceptance.","tokens_in":11975,"tokens_out":27918,"duration_ms":287046,"concrete_test":"Obtain [16] and independently re-derive Eq. (3.30) of Lemma 3.22 for the class K^α_FNO(K) defined in equation (9) instead of U^{α,∞}_{ℓ,NO}; specifically, check that the covering-number/entropy estimate used in the proof of [16, Thm. 3.21] holds with constants independent of m and with the stated dependence on α. If the entropy estimate cannot be re-proved for K^α_FNO, then the upper bound β*≤1/2 in Thm. 5 does not follow.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that architecture expressivity alone does not produce faster-than-n^{-1/2} rates depends on Thm. 5's upper bound β*≤1/2. The proof note states that Eq. (3.30) of [16, Lem. 3.22] 'remains valid' with K^α_FNO in place of U^{α,∞}_{ℓ,NO}, based on a 'perusal', and then cites [16, Eqn. 2.6] and the proof of [16, Thm. 3.21]. No details are given for why the metric-entropy or approximation estimates transfer to the class defined in (9). If the two classes differ in a way that breaks the argument, the 1/2 ceiling is unsupported and the 'regularity not expressivity' conclusion loses its FNO half. This is not an observed mathematical error, but it is a missing verification at a load-bearing joint of the survey.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey paper presents a unified, notationally consistent overview of recent results in operator learning theory with emphasis on sample complexity. Section 2 reviews two empirical-risk-minimization bounds for holomorphic operators: Theorem 1, adapted from Reinhardt–Wang–Zech, gives an expectation error bound with a rate approaching n^{-1/2} for deep ReLU networks; Theorem 2, adapted from Adcock–Dexter–Moraga, gives a faster-than-Monte-Carlo algebraic rate for a handcrafted tanh-network class under polynomial truncation-error decay and bounded noise. Section 3 surveys minimax results: Theorem 3 gives a polylogarithmic lower bound for C^k and Lipschitz operator classes from point evaluations; Theorem 4 gives lower bounds for holomorphic classes showing rates of order n^{-(1/p-1/2)}; Theorem 5 gives two-sided bounds for FNO-approximable classes, with optimal algebraic exponent between roughly 1/2 and 1/2; and Theorem 6 extends the story to a noisy sampling setting. The paper closes with several open problems. The authors are transparent that they are re-presenting results from cited papers in a unified notation, with a partial loss of generality.","tokens_in":12282,"tokens_out":17155,"duration_ms":163481,"significance":"If the results are correct as presented, the survey provides a valuable synthesis: holomorphy yields arbitrarily fast algebraic minimax rates, Lipschitz/C^k regularity yields only polylogarithmic rates, and FNO-approximability yields algebraic but no-better-than-n^{-1/2} rates. This gives a clear organizing principle, essentially 'regularity governs sample complexity.' The explicit attribution of each theorem to a cited source is a strength, as is the careful separation of upper and lower bounds and the candid discussion of open problems. The paper does not prove new theorems, but a survey does not need to; its contribution is the coherent framing. The weakest point is a single load-bearing verification in Theorem 5, discussed below; if that is repaired, the survey would be a useful reference for the operator-learning community.","major_comments":[{"comment":"The upper bound β*≤1/2 is load-bearing for the FNO half of the paper's central conclusion, yet the proof rests on the sentence that 'perusal' of the proof of [16, Lem. 3.22] shows that [16, Eqn. (3.30)] remains valid with K^α_FNO(K) in place of U^{α,∞}_{ℓ,NO}. This is not a formality: the class K^α_FNO(K) defined in Eq. (9) involves a specific FNO architecture, parameter norm bound exp(m), and sup_{m} m^α e_K(f,N_m^{FNO}) approximation error, while the class in [16] may measure approximation differently. The paper gives no argument that the metric-entropy lower bound or the relevant construction transfers. Please supply a proof or a precise reduction—for example, show that U^{α,∞}_{ℓ,NO} embeds into K^α_FNO(K) with explicit parameter choices, or state and prove a lemma establishing the lower bound directly for K^α_FNO(K). Without this, Eq. (10)'s upper bound is unsupported and the 'no fa","section":"§3, Theorem 5 and the proof paragraph after Eq. (10)"},{"comment":"The displayed theorem contains only lower bounds on s_n(H(b)∘ι), but the text states that the rate is 'optimal, up to log factors.' This optimality conclusion is not a consequence of the theorem as stated; it requires a matching upper bound, which is only indirectly available through Theorem 2 under extra hypotheses (e.g., Eq. (6), σ=0). Please state the matching upper bound explicitly, or rephrase the optimality claim to identify precisely which theorem supplies the matching upper bound and under which assumptions. This would prevent a reader from attributing to Theorem 4 an assertion it does not contain.","section":"§3, Theorem 4 and the paragraph following it"}],"minor_comments":[{"comment":"Both theorems are stated for global minimizers of the nonconvex ERM objective (3). Existence is asserted or guaranteed with high probability, but no algorithm is claimed to find such minimizers. The paper acknowledges this only in the §4 open-problem discussion. Since the abstract promises convergence rates for empirical risk minimization, add a remark in §2 clarifying the oracle nature of these results and referring the reader to the open problems.","section":"§2, Theorems 1–2"},{"comment":"The notation e_K(f,𝒩) in the definition of Γ^α_FNO is ambiguous because 𝒩 is not defined at that point; it should be e_K(f,𝒩_m^{FNO}). Also, the rendering of the map ι as U→R^N should be U→ℝ^ℕ, and the inclusion H_{0,1}=[−1,1]^ℕ⊆R(b) is meant in ℓ^∞, not ℓ^2; clarify the ambient space.","section":"§3, Eq. (9)"},{"comment":"The displayed chain contains three inequalities, but the following text refers only to 'the first' and 'the second' of them. Label the inequalities (e.g., (12a), (12b), (12c)) to make the explanation unambiguous.","section":"§3, Theorem 6, Eq. (12)"},{"comment":"The additive τ in the right-hand side means that for fixed τ>0 the displayed bound does not tend to zero as n→∞. The text's phrase 'approximate Monte Carlo rate' is intended to address this, but the dependence of c on τ is not stated. Please write c_τ or add a sentence clarifying that c depends on τ and other fixed parameters, and that the τ term is to be sent to zero in the limiting-rate statement.","section":"§2, Theorem 1, Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written and useful survey, and I am sympathetic to accepting it. The only real obstacle is the unsupported transfer in the proof of Theorem 5's upper bound; that is a load-bearing point, but it is likely fixable within the manuscript's scope by adding a short proof or a precise reduction to the cited lemma. If the authors supply that missing verification, I would be willing to accept. The other issues are presentation-level and do not affect my recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nHere's the honest take: this is a survey, not a research paper. Its six numbered theorems are explicitly corollaries of results from other groups, and the authors say so upfront. The novelty is low, but the synthesis is valuable: the paper gives a clean, unified notation for ERM bounds and minimax rates across holomorphic, Lipschitz, and FNO-approximable operator classes, with careful attribution and explicit statements about what was simplified. I agree with the reader that this deserves an ACCEPT verdict as a survey, with the understanding that it's not a primary research contribution.\n\nThe strongest parts are the minimax section. The polylogarithmic barrier for Lipschitz/C^k operators (Thm. 3) and the fast algebraic rates for holomorphic classes (Thm. 4) are presented clearly. The discussion of FNO-approximable classes in Thm. 5, with the exponent between 1/(2+8/α) and 1/2, is a helpful way to frame the 'regularity, not expressivity' takeaway. The open problems are sensible and not padded.\n\nThe main soft spot is in Thm. 5. The upper bound β*≤1/2 relies on the assertion that [16, Lem. 3.22]'s inequality 'remains valid' with K^α_FNO in place of U^{α,∞}_{ℓ,NO}, based on a 'perusal' of the proof. That transfer is not shown. This is a load-bearing step: if the two classes differ in the details, the 1/2 ceiling could fail, and with it the FNO half of the 'regularity, not expressivity' story. This is not an observed error, but it is a missing verification at a critical joint. A referee should ask for the details, or for the statement to be softened to a conjecture.\n\nAlso, the ERM upper bounds in §2 are existential for global minimizers of a nonconvex objective. The paper notes this only in the conclusion; it's a real caveat but not a fatal flaw, since the same issue applies to most neural-network theory.\n\nWho this is for: graduate students or researchers looking for a map of the sample-complexity landscape in operator learning. It's a useful reference, and I'd be comfortable citing it as a review. It deserves a serious referee, but the referee should focus on the Thm. 5 transfer and the algorithmic caveat.\n\nBest,\n[You]","headline":"A transparent survey of operator learning sample-complexity bounds; the FNO exponent ceiling in Thm. 5 rests on an unverified lemma transfer.","tokens_in":12834,"tokens_out":3480,"would_cite":true,"duration_ms":37498,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["41A46","47A58","62C20","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the sample complexity of learning an operator from point evaluations is governed by the operator's regularity, proving that holomorphic operators admit algebraic error rates whereas Lipschitz or C^k operators cannot.","keywords":["operator learning","empirical risk minimization","minimax rates","holomorphic operators","neural networks","Fourier Neural Operators","n-widths","sample complexity"],"falsifier":"Find a concrete learning method—neural or otherwise—that achieves an algebraic error rate for the unit Lipschitz ball from n point evaluations, contradicting Theorem 3; or construct a holomorphic class with l^p-summable parameters whose minimax rate is strictly slower than n^{-(1/p-1/2)}, contradicting Theorem 4.","tokens_in":11885,"feed_emoji":"📈","tokens_out":9767,"duration_ms":88537,"temperature":0.7,"pith_summary":"This survey of operator learning theory sets out to clarify how many samples are needed to approximate a nonlinear map between function spaces from point evaluations. Drawing together recent upper bounds for empirical risk minimization and lower bounds from minimax analysis, it proposes that the smoothness of the target operator—not the expressivity of the neural network—is the decisive factor. The paper establishes that holomorphic operators can be learned at error rates arbitrarily close to the Monte Carlo rate n^{-1/2} (or faster), while merely Lipschitz or C^k operators suffer a 'curse of sample complexity' with only polylogarithmic decay. Fourier Neural Operator–approximable classes sit in between, with optimal algebraic rate exponents between 1/(2+16/α) and 1/2. If these results hold, they tell practitioners where to place their bets: analytic regularity justifies optimism, while Lipschitz-only problems are fundamentally hard.","feed_headline":"Holomorphic operators beat the sample-complexity curse","feed_subtitle":"A survey of minimax bounds shows operator regularity, not architecture, decides how fast neural operators learn from data.","key_machinery":"The central objects are the nonlinear sampling n-width s_n(K)_X, which records the best worst-case error achievable with n point evaluations over a class K, and the regularity classes themselves: the holomorphic classes H(b) built from Bernstein polyellipses with l^p-summability, and the FNO-approximable classes K_alpha^FNO. The upper bounds are obtained by balancing approximation error from finite-dimensional encoding/decoding with statistical error from finite samples, using empirical process theory (for ReLU networks) and compressed sensing (for handcrafted tanh networks). The lower bounds combine n-width arguments with concrete constructions of hard operator families. The key mechanism i","core_discovery":"On the paper's own terms, the central discovery is a hierarchy of minimax learnability for operator classes. For the unit ball in the Lipschitz or C^k spaces, the best achievable worst-case error over all methods using n point evaluations decays only polylogarithmically, so no algebraic rate is possible (Thm. 3). For holomorphic operators with l^p-summable parameter sequences, the minimax error is n^{-(1/p-1/2)} up to log factors, a rate that can be made arbitrarily close to n^{-1/2} as p→0 and is attained by an empirical risk minimizer based on compressed sensing (Thms. 2 and 4). For operators that are efficiently approximated by Fourier Neural Operators at rate α, the optimal algebraic min","pith_inferences":["A practical consequence of the separation is that estimating the analytic regularity of a solution map (e.g., from parameter-to-solution maps of PDEs) before choosing an architecture would be more valuable than architecture search; holomorphy justifies expecting fast convergence.","The lower bounds assume point-evaluation measurements; if linear functionals (e.g., Fourier coefficients) are allowed as measurements, the Lipschitz lower bound may not apply, and the sample complexity could change—this is not explored in the survey.","The theorem results are existential for global minimizers of a nonconvex problem; without proofs that gradient-based training can reach these minimizers, the theoretical rates may not be realized in practice, an open problem the paper acknowledges.","One could test the FNO upper bound by training FNOs on holomorphic operators: if the observed error exponent exceeds 1/2 for large α, the true minimax exponent might be higher than the survey's guess."],"forward_implications":["For holomorphic operators with l^p-summable parameter sequences and negligible encoder/decoder error, empirical risk minimization achieves minimax-optimal error rates n^{-(1/p - 1/2)} in the absence of noise, which are algebraic and can be made arbitrarily fast as p decreases (Thms. 2 and 4).","No learning method can achieve algebraic sample complexity for the unit Lipschitz or C^k ball from point evaluations; the best possible worst-case error decays only polylogarithmically (Thm. 3).","For operators that are well approximated by Fourier Neural Operators at rate α, the optimal algebraic minimax rate exponent is between 1/(2+16/α) and 1/2, so algebraic rates are possible but capped at n^{-1/2} (Thm. 5).","The minimax errors for the holomorphic and FNO classes are both bounded above by the minimax error for the Lipschitz class, giving a clean separation between regularity classes (paper's §3 discussion)."],"fun_headline_variants":["Holomorphic operators achieve near-optimal minimax rates","Operator learning: holomorphy gives near-optimal minimax rates","Smoothness alone won't get you algebraic rates in operator learning","Holomorphic operators: sample complexity near n^-1/2","Minimax bounds for operator learning: holomorphy beats smoothness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the empirical risk minimization problem (3) has a solution and that training can find it: the theorems are existential for global minimizers, and no algorithm is proven to reach them, so the sample-complexity rates apply only to an idealized optimization.","fun_headline_variants_meta":{"raw":{"variants":["Holomorphic operators achieve near-optimal minimax rates","Operator learning: holomorphy gives near-optimal minimax rates","Smoothness alone won't get you algebraic rates in operator learning","Holomorphic operators: sample complexity near n^-1/2","Minimax bounds for operator learning: holomorphy beats smoothness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000813,"raw_usage":{"total_tokens":3335,"prompt_tokens":613,"completion_tokens":2722,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":357,"completion_tokens_details":{"reasoning_tokens":2634}},"tokens_in":357,"tokens_out":2722,"duration_ms":21928,"temperature":1.0,"reasoning_tokens":2634,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T19:47:39.819213+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a concrete learning method—neural or otherwise—that achieves an algebraic error rate for the unit Lipschitz ball from n point evaluations, contradicting Theorem 3; or construct a holomorphic class with l^p-summable parameters whose minimax rate is strictly slower than n^{-(1/p-1/2)}, contradicting Theorem 4.","supporting_citations":[],"review_version":1}