{"id":"e94d6f0b-d2b9-4209-b9f3-00a18e9f89d7","arxiv_id":"2412.08453","paper_version":3,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Sums of ell-variate ridge functions approximate r-smooth Sobolev functions on R^d at the optimal rate n^{-r/(d-ell)}, giving sharp rates for generalized translation networks and complex-valued neural networks.","lead":"This paper proves sharp approximation rates for Sobolev functions by sums of multivariate ridge functions, establishing an optimal error of n^{-r/(d-ell)}. The result resolves a natural multivariate analog of known univariate bounds and yields optimal rates for generalized translation networks and complex-valued neural networks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniform L1 boundedness of the quasi-projections in Prop. 2.7(3) is the linchpin; it rests on [4, Thm 11.4.1] at a parameter/hypothesis boundary not verified in the text.","rationale":"The reader's weakest assumption identifies the same step: Proposition 2.7(3) is the only place where the proof needs an estimate that is uniform in the polynomial degree s, and Lemma 3.8 depends on it directly. I agree that this is the most load-bearing point in the proof of the lower bound. The rest of the argument appears internally coherent: Lemma 3.5 correctly separates the inner product of a ridge function and a polynomial into finitely many polynomial coefficients; the sign-set cardinality argument follows [17] with the required degree and dimension bookkeeping; the upper bound via polynomial ridge functions is clean; and the GTN/CVNN applications are legitimate consequences of Theorems 1.1 and 1.2. The cited Cesaro-mean theorem is a standard tool in harmonic analysis on the ball, and the explicit parameter choice (kappa with last coordinate 1/2, integer sigma > d/2) is consistent with the known form of such results. However, the paper does not verify the theorem's exact hypotheses for p = 1, and because the companion proposition in [4] is stated without proof there, the step deserves an explicit check. Since no internal inconsistency is apparent and the concern is about verification of a standard cited result rather than a demonstrated error, I do not move the reader's ACCEPT verdict; the proposed check is a prudent confirmation of the linchpin.","tokens_in":50201,"tokens_out":26778,"duration_ms":273953,"concrete_test":"Pull up [4, Theorem 11.4.1] and verify: (a) the admissible p-range includes p = 1, (b) with kappa = (0,...,0,1/2) the critical index is strictly less than sigma for every integer sigma > d/2, and (c) the implied constant is independent of k and s. If the book is unavailable, run a numerical check for d = 2, sigma = 2: compute ||S_k^sigma||_{L1(B_2)->L1(B_2)} for k up to, say, 1000 using the Section 2.3 basis; uniform boundedness supports Prop. 2.7(3), while growth in k would refute it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 2.7(3) is used exactly once at the critical transfer Lemma 3.8: without sup_s ||Pr_s||_{L1(Bd)->L1(Bd)} < infinity, the lower bound for Pr_s(R*_{n,d,ell}) cannot be pushed down to R*_{n,d,ell}, and Theorem 1.1 fails. The proof of Prop 2.7(3) is otherwise elementary (Lemma 2.6, finite-difference estimates, binomial bounds), but the decisive uniform-in-k estimate ||S^sigma_k(f)||_{L1} <= C3(d,sigma)||f||_{L1} is imported from [4, Theorem 11.4.1]. The paper chooses kappa = (0,...,0,1/2) in R^{d+1} and sigma > d/2, but does not state the theorem's exact critical-index assumptions for p = 1. Since [4] itself does not prove the companion proposition (as the paper notes), this is a genuine external dependency at the core of the central claim. If the true condition were, for instance, sigma > d/2 + kappa_{d+1}, or if the theorem only covers p > 1, the constant C3 would not exist and Lemma 3.8 would collapse.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies approximation of Sobolev functions on the unit ball B_d by sums of n multivariate ridge functions x ↦ ρ(Ax) with A ∈ R^{ℓ×d}. The main results (Theorems 1.1 and 1.2) establish matching lower and upper bounds of order n^{-r/(d-ℓ)} for the worst-case L^q error over the unit ball of W^{r,p}(B_d), under the assumptions that the lower bound holds for all p,q∈[1,∞] and the upper bound for 1≤q≤p≤∞. The lower-bound proof uses a volumetric/sign-pattern argument together with a family of quasi-projection operators P_r^s that are uniformly bounded on L^1, and the upper bound is obtained by expressing polynomials as sums of ridge polynomials with fixed matrices. The paper also applies these results to generalized translation networks and complex-valued neural networks, obtaining optimal rates n^{-r/(d-ℓ)} and n^{-r/(2d-2)}, and it identifies and repairs a gap in the earlier proof for univariate ridge functions in [18].","tokens_in":50502,"tokens_out":9719,"duration_ms":93442,"significance":"If the external hypotheses on which the proof depends are verified, this paper solves the open problem of sharp approximation rates for multivariate ridge functions and provides the first optimal rates for shallow generalized translation networks and complex-valued neural networks with general activation functions. The paper is careful and detailed, with explicit constants, and its replacement of the orthogonal projection in [18] by uniformly bounded quasi-projections is a genuinely new methodological step that also corrects a known gap. The results are sharp and falsifiable, and the applications to neural networks are concrete. The principal caveat is that the central lower bound rests on an imported uniform-L^1 boundedness estimate for Cesàro means, whose hypotheses are not fully checked in the manuscript.","major_comments":[{"comment":"The uniform L^1 bound for the quasi-projections P_r^s is the linchpin of the lower bound. The proof of Proposition 2.7(3) derives the estimate ∥P_r^s f∥_{L^1} ≤ C∥f∥_{L^1} from the uniform-in-k bound ∥S^σ_k(f)∥_{L^1} ≤ C_3∥f∥_{L^1}, which is imported from [4, Theorem 11.4.1] with the parameter choice κ=(0,…,0,1/2)∈R^{d+1} and σ>d/2. The manuscript does not state the hypotheses of that theorem, and it explicitly notes that the corresponding proposition is stated without proof in [4]. Since Lemma 3.8 transfers the lower bound from P_r^s(R*_{n,d,ℓ}) to R*_{n,d,ℓ}, the whole lower bound of Theorem 1.1 collapses if this estimate fails (for instance, if the theorem requires σ>d/2+κ_{d+1} or excludes p=1). The authors should either provide a self-contained proof of the uniform L^1 bound or quote [4, Theorem 11.4.1] in full and verify that the chosen parameters satisfy all of its assumptions.","section":"Section 2.3, Proposition 2.7(3) and Lemma 3.8"},{"comment":"The sign-set cardinality bound in Lemma 3.1, which is essential for the entropy lower bound, relies on Lemma 3.3, imported verbatim from [17, Lemma 3] without proof. Similarly, the upper bound in Section 4 depends on [26, Corollary 5.12]. These are external results at load-bearing steps of the argument. Given that the paper's contribution is precisely a careful and self-contained repair of a gap in [18], the reader needs to be able to verify that these external lemmas apply. The authors should either include proofs of these lemmas or state them with all hypotheses and exact reference locations, so that no hidden condition (e.g., on the degree bounds or on N+K) is left unchecked.","section":"Section 3.2, Lemma 3.1 and Lemma 3.3"}],"minor_comments":[{"comment":"In the displayed chain of inequalities after equation (3.31), the factor should be (2ϑ)^{-r}, not (2ϑ)^r. The subsequent conclusion ∥f_{ε*}−P∥_{L^1} ≥ a·c_2·c_3·2^{-r}·m^{-r/d} is consistent with the negative exponent, so this appears to be a typo rather than a mathematical error.","section":"Section 3.3, Proposition 3.7"},{"comment":"The constant C(d) is written as C = C_1C_2C_3·2^{σ+1}, but C_3 depends on d and σ, and σ is chosen depending on d; this is fine, but the notation could be made clearer by writing C = C(d,σ) and then noting that σ is fixed by d.","section":"Section 2.3, proof of Proposition 2.7(3)"},{"comment":"The equation numbering in the discussion of the gap in [18] is slightly confusing: the manuscript refers to 'Equation (14)' of [18] but then labels its own displayed equation as (A.4). This could be clarified by explicitly saying that (A.4) is the corresponding statement in the notation of the present paper.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the journal and the main results are likely to be influential. The main risk is not internal inconsistency but the external dependency on [4, Theorem 11.4.1] for the uniform L^1 bound of the quasi-projections. If the authors can verify the hypotheses or supply a proof, I would be happy to support acceptance. The remaining external dependencies ([17, Lemma 3], [26, Corollary 5.12]) are standard but should also be made explicit for completeness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a strong paper. It proves the first sharp n^{-r/(d-\\ell)} approximation rates for sums of \\ell-variate ridge functions on Sobolev classes, and the applications to generalized translation networks and complex-valued networks are genuine corollaries rather than afterthoughts. The CVNN rate n^{-r/(2d-2)} versus n^{-r/(2d-1)} for real networks is substantive. The paper also identifies and repairs a real gap in Maiorov's L^p ridge-function lower bound, and the appendix showing why the orthogonal-projection argument fails is useful on its own.\n\nWhat I like: the proof structure is clear, the quasi-projection substitution is the right remedy, and the authors are unusually candid about where the argument depends on external results. The lower bound holds under weak integrability assumptions, and the upper bound uses only fixed polynomial ridge functions, which is clean.\n\nSoft spots, in order. First, the stress-test concern about Proposition 2.7(3) is legitimate. The uniform L1 boundedness of the quasi-projections is the linchpin of Lemma 3.8, and the decisive Cesaro-mean estimate is imported from Dai-Xu [4, Thm 11.4.1] with a parameter choice (kappa = (0,...,0,1/2), sigma > d/2) whose hypotheses are not stated in the text. Since [4] itself omits the proof of the relevant companion proposition, this is a genuine external dependency at the core of the lower bound. I do not know whether the theorem actually requires sigma > d/2 + kappa_{d+1} or p > 1; if it does, the bound as written collapses. A referee should ask the authors to state the exact theorem and verify the parameter ranges. This is likely fixable, but it is load-bearing, not cosmetic.\n\nSecond, the rank restriction in Theorem 1.3(1) for L1_loc activations is left open; that is minor and honestly flagged. Third, the argument is long and not machine-checked; the main structural steps are clear, so I do not see that as a serious weakness.\n\nWho this is for: approximation theorists and anyone working on neural network expressivity rates. It deserves a serious referee and likely acceptance after the external hypotheses are checked.","headline":"Sharp multivariate ridge rates with an honest fix of Maiorov's gap; the only real risk is an imported Cesaro-boundedness theorem at the core of the lower bound.","tokens_in":51056,"tokens_out":1927,"would_cite":true,"duration_ms":21292,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["41A30","41A25","41A63","46E35","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes that n sums of ℓ-variate ridge functions approximate r-times-differentiable functions on the ball of R^d with sharp error n^{-r/(d-ℓ)}, and transfers this order to generalized translation and complex-valued networks.","keywords":["multivariate ridge functions","Sobolev functions","optimal approximation rates","generalized translation networks","complex-valued neural networks","quasi-projection operators","sign pattern counting","sharp bounds"],"falsifier":"Compute sup_s ||Pr_s||_{$L^{1}$(B_d)→$L^{1}$(B_d)} for the quasi-projection of Section 2.3: if the norms are unbounded in s, the transfer lemma (Lemma 3.8) fails and with it the lower bound; conversely, an explicit Sobolev function of regularity r approximated by n ℓ-variate ridge sums at error much smaller than $n^{{-r/(d-ℓ)}}$ would refute Theorem 1.1 directly.","tokens_in":50030,"feed_emoji":"📉","tokens_out":8250,"duration_ms":79220,"temperature":0.7,"pith_summary":"This paper proves a sharp approximation theorem: any L^p-Sobolev function of regularity r on the unit ball of R^d is approximated, up to error of order $n^{{-r/(d-ℓ)}}$, by a sum of n ridge functions of ℓ variables, and no such sum can do asymptotically better. The lower bound holds even in the hardest regime, approximating L^∞-Sobolev targets with error measured in $L^{1}$, while the upper bound is achieved using only polynomial ridge functions with fixed matrices. This settles the multivariate generalization of the classical univariate ridge-function results. As consequences, generalized translation networks with activation dimension ℓ attain exactly $n^{{-r/(d-ℓ)}}$, improving as ℓ grows, and a constructed smooth complex activation function lets complex-valued shallow networks reach $n^{{-r/(2d-2)}}$, better than the real-network benchmark $n^{{-r/(2d-1)}}$.","feed_headline":"Ridge sums hit the sharp rate n^{-r/(d-ℓ)}","feed_subtitle":"The same exponent governs translation networks and shows complex nets beat real ones.","key_machinery":"The argument rides on a quasi-projection operator Pr_s: $L^{1}$(B_d) → P_{2s-1}(B_d) that fixes polynomials of degree up to s and, unlike the orthogonal projection, is uniformly bounded in $L^{1}$ for all s (Proposition 2.7). This operator transfers the lower bound from the projected ridge class Pr_s(R*_{n,d,ℓ}) to the actual class R*_{n,d,ℓ} through Lemma 3.8. The lower bound itself is built from sign-set counting: after separating variables in the inner product ⟨ρ(A·), P⟩ (Lemma 3.5), the sign patterns arising from projected ridge functions are few enough (Lemma 3.1) that a standard volume argument (Lemma 3.2) finds a Sobolev bump function far from every such pattern.","core_discovery":"The central discovery is that the approximation order of Sobolev functions by sums of ℓ-variate ridge functions is governed only by the co-dimension d−ℓ of the ridge subspace: an asymptotic error of order $n^{{-r/(d-ℓ)}}$ is simultaneously necessary and sufficient. The lower bound holds for every p,q ∈ [1,∞], the matching upper bound for 1 ≤ q ≤ p ≤ ∞, and the upper bound is attainable with polynomial ridge functions and a fixed choice of matrices. The proof also closes a gap in the univariate case by replacing the orthogonal projection onto the polynomials of degree at most s, which is not uniformly bounded in $L^{1}$, with a quasi-projection onto P_{2s-1}(B_d) that is uniformly $L^{1}$-bounded. The same rate then transfers to generalized translation networks and, through the Wirtinger calculus, to complex-valued neural networks.","pith_inferences":["The exponent −r/(d−ℓ) suggests a co-dimension law: the approximation difficulty in R^d is set by the dimension of the space orthogonal to the ridge directions, and one could test whether other structured model classes (tensor trains, subspace networks, dictionary models) obey an analogous law with their own intrinsic co-dimension.","Because only the Jackson inequality enters the upper bound, the n^{-r/(d-ℓ)} rate plausibly extends to Besov or other smoothness scales with the same regularity parameter r, a direct extrapolation of Remark 4.3.","The two-step lower-bound scheme (cardinality bound for sign patterns, then a volume argument on separated cubes) is transferable and should yield sharp rates for other constrained dictionaries whose projected parameter sets have small dimension.","One could empirically test the constructed piecewise activation for small dimensions (for instance d=3, ℓ=2, r=2, where the predicted rate is n^{-2}): a materially faster decay would point to a hidden structural assumption, while a slower decay would suggest the constant prefactor is large."],"forward_implications":["Generalized translation networks with activation dimension ℓ approximate L^p-Sobolev functions of regularity r at the sharp rate n^{-r/(d-ℓ)}, so raising ℓ strictly improves the achievable rate (Theorem 1.3).","For each ℓ there is a single smooth activation function τ: R^ℓ → R for which the optimal rate is achieved with fixed weight matrices, whatever d, p, q and r (Theorem 1.3(2)).","Shallow complex-valued networks attain n^{-r/(2d-2)}, beating the real benchmark n^{-r/(2d-1)} obtained by identifying C^d with R^{2d} (Theorem 1.4).","The L^1-norm lower bound repairs the gap in the earlier univariate proof by replacing orthogonal projection with a uniformly L^1-bounded quasi-projection (Appendix A).","The upper bound needs no smoothness beyond a Jackson-type polynomial approximation inequality, so the same rate holds for any function class satisfying Proposition 2.5 (Remark 4.3)."],"supporting_citations":[{"why":"Lays down the univariate lower-bound scheme (sign sets, separated bumps) that this paper generalizes to ℓ > 1, and whose Lemma 5 is shown in Appendix A to contain a gap that the quasi-projection fixes.","marker":"[18]"},{"why":"Supplies Lemma 3 (sign-set cardinality bound) and Theorem 3 (inner-product separation), both of which are generalized here to multivariate ridge functions.","marker":"[17]"},{"why":"Provides the Cesàro-means boundedness theorem and the quasi-projection framework that yield the uniform L^1 bound of Proposition 2.7.","marker":"[4]"},{"why":"Section 5 supplies the polynomial-ridge representation used to prove the upper bound (Proposition 4.1) and its complex-counterpart Theorem 5.7.","marker":"[26]"},{"why":"Provides the Jackson-type inequality (Proposition 2.5) used in the upper bound, and the continuous-weight-selection lower bound that this paper's GTN result explicitly contrasts with.","marker":"[21]"},{"why":"Supplies the bespoke sigmoidal activation function with which every ridge sum can be realized as a shallow network, used in the GTN upper-bound construction.","marker":"[19]"},{"why":"Constructs the complex activation function (Lemma F.4) and the earlier n^{-r/(2d-1)} CVNN rate that Theorem 1.4 improves to n^{-r/(2d-2)}.","marker":"[9]"}],"fun_headline_variants":["Co-dimension alone sets ridge approximation rate for Sobolev","Exact ridge error: n^-r/(d-ℓ) for all Sobolev spaces","Sharp ridge bounds extend to translation and complex nets","Ridge functions: co-dimension determines approximation order","Complex nets outperform real ones via ridge sharp rate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The lower bound rests on the quasi-projection operators being uniformly bounded in the $L^{1}$ norm for every polynomial degree, a fact whose proof invokes a Cesàro-mean theorem from a cited textbook.","fun_headline_variants_meta":{"raw":{"variants":["Co-dimension alone sets ridge approximation rate for Sobolev","Exact ridge error: n^-r/(d-ℓ) for all Sobolev spaces","Sharp ridge bounds extend to translation and complex nets","Ridge functions: co-dimension determines approximation order","Complex nets outperform real ones via ridge sharp rate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000411,"raw_usage":{"total_tokens":2132,"prompt_tokens":954,"completion_tokens":1178,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":1093}},"tokens_in":570,"tokens_out":1178,"duration_ms":11607,"temperature":1.0,"reasoning_tokens":1093,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:45:17.475608+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute sup_s ||Pr_s||_{$L^{1}$(B_d)→$L^{1}$(B_d)} for the quasi-projection of Section 2.3: if the norms are unbounded in s, the transfer lemma (Lemma 3.8) fails and with it the lower bound; conversely, an explicit Sobolev function of regularity r approximated by n ℓ-variate ridge sums at error much smaller than $n^{{-r/(d-ℓ)}}$ would refute Theorem 1.1 directly.","supporting_citations":[{"cited_title":"Best approximation by ridge functions inLp-spaces","cited_arxiv_id":null,"evidence_quote":"Lays down the univariate lower-bound scheme (sign sets, separated bumps) that this paper generalizes to ℓ > 1, and whose Lemma 5 is shown in Appendix A to contain a gap that the quasi-projection fixes."},{"cited_title":"On Best Approximation by Ridge Functions","cited_arxiv_id":null,"evidence_quote":"Supplies Lemma 3 (sign-set cardinality bound) and Theorem 3 (inner-product separation), both of which are generalized here to multivariate ridge functions."},{"cited_title":"Dai and Y","cited_arxiv_id":null,"evidence_quote":"Provides the Cesàro-means boundedness theorem and the quasi-projection framework that yield the uniform L^1 bound of Proposition 2.7."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Section 5 supplies the polynomial-ridge representation used to prove the upper bound (Proposition 4.1) and its complex-counterpart Theorem 5.7."},{"cited_title":"Neural networks for optimal approximation of smooth and analytic functions","cited_arxiv_id":null,"evidence_quote":"Provides the Jackson-type inequality (Proposition 2.5) used in the upper bound, and the continuous-weight-selection lower bound that this paper's GTN result explicitly contrasts with."},{"cited_title":"Lower bounds for approximation by MLP neural networks","cited_arxiv_id":null,"evidence_quote":"Supplies the bespoke sigmoidal activation function with which every ridge sum can be realized as a shallow network, used in the GTN upper-bound construction."},{"cited_title":"Optimal approximation using complex-valued neural networks","cited_arxiv_id":null,"evidence_quote":"Constructs the complex activation function (Lemma F.4) and the earlier n^{-r/(2d-1)} CVNN rate that Theorem 1.4 improves to n^{-r/(2d-2)}."}],"review_version":1}