{"id":"fc8a50bc-e5f0-4f8c-a607-8c82f5c8d43f","arxiv_id":"2507.06789","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ReLU networks approximate spectral Barron functions of smoothness as low as 1/2 at the N^{-1/2} rate, and L-layer networks achieve sharp N^{-sL} rates for 0<sL<=1/2.","lead":"This paper proves that neural networks approximate rough spectral Barron functions at the optimal Monte Carlo rate: smoothness s=1/2 suffices for a shallow network, and depth L gives rate N^{-sL} for 0<sL<=1/2. A sharp lower bound shows these rates cannot be improved except by a logarithmic factor in the uniform norm.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection to the central theorems; Lemma 3.2 holds, so the reader's weakest assumption does not land; remaining issues are editorial.","rationale":"I traced the central argument in good faith. Lemma 3.2, which the reader identified as the weakest assumption, is actually a standard period-doubling identity for symmetric functions and holds for the specific gamma(t,r) defined in Section 3.1. The proof of Theorem 3.1 correctly constructs gamma_{,n_xi} via repeated composition, keeps t_xi(x) in [0,1], counts ReLU units per layer, and combines Markov bounds to get the N^{-sL} rate with a sqrt(p) prefactor. Theorem 3.6's Rademacher-complexity proof, including the covering-number estimate and the logarithmic factor, is plausible and internally consistent; the constants are loose but the rate is right. The lower bound Theorem 3.10 uses a cosine with 2^{L+2}N^L oscillations and the standard bound on sign changes of ReLU networks, giving a matching N^{-sL} lower bound up to a constant. I found no flaw that would change the reader's conditional verdict. The reader's specific weakest assumption does not land, hence 'disagree' on that point. The genuine concerns are the abstract's overclaim of dimension-free prefactors and omitted derivations in corollaries; these are real editorial issues but do not affect the correctness of Theorems 3.1, 3.6, and 3.10, so keeping the CONDITIONAL verdict unchanged is appropriate.","tokens_in":24113,"tokens_out":45108,"duration_ms":422975,"concrete_test":"Independently verify Lemma 3.2 for the paper's gamma(t,r) by direct computation: evaluate (gamma_{,n2} ∘ beta_{,n1})(t) and gamma_{,2 n1 n2}(t) at t = j/(2 n1 n2) and t = (2j+1)/(4 n1 n2) for n1=1,n2=1 and n1=2,n2=3, checking left and right limits at breakpoints; if the values agree, the composition identity and hence the per-layer width bound 4 ceil((1+|xi|_1)^{1/L}) are sound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection to the central claim identified. The reader's flagged weakest assumption, Lemma 3.2, is the composition identity gamma_{,n2} ∘ beta_{,n1} = gamma_{,2 n1 n2}, quoted from [20]. Direct verification for the specific gamma(t,r) used here shows the identity is sound: on each interval of beta_{,n1}, the argument traverses [0,1] once upward and once downward, and gamma's symmetry about 1/2 makes the downward traversal coincide with the upward one, producing 2 n1 n2 periods. The affine substitution t_xi(x) lies in [0,1] because n_xi ≥ 1+|xi|_1, so the identity applies. The Lp proof of Theorem 3.1 and the L∞ proof of Theorem 3.6 are internally consistent, and the lower bound Theorem 3.10 uses standard oscillation counting via [39, Lemma 3.2]. Secondary, revision-level issues remain: the abstract's 'dimension-free prefactors' conflicts with the sqrt(1+dL ln N) prefactor in Theorem 3.6 and Corollary 3.9, and Corollaries 3.4, 3.9, and the final Lp improvement omit derivations. These do not undermine the main theorems' correctness.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proves non-asymptotic approximation bounds for spectral Barron functions by shallow sigmoidal and deep ReLU networks. For f in B^s(Omega) with 0 < s <= 1/2 and p >= 2, Theorem 2.4 gives a shallow sigmoidal network with N units satisfying ||f - f_N||_{Lp} <= 28 pi sqrt(p) ||f||_{B^s} N^{-s}. Theorem 3.1 extends the sharp rate N^{-sL} to (L,N)-ReLU networks under 0 < sL <= 1/2, with constant 13 sqrt(p). Theorem 3.6 establishes the uniform bound ||f - f_{L,N}||_{Linfty} <= 45 ||f||_{B^s} N^{-sL} sqrt(1 + dL ln N). Theorem 3.10 provides a matching L^p lower bound for 1 <= p <= infinity, showing the rates are sharp up to a logarithmic factor. Corollaries refine the bounds to depend on the seminorm upsilon_{f,s}; Appendix A proves a sharp bound on the Khintchine constant, and Appendix B shows the embedding B^s into C^s is essentially sharp.","tokens_in":24330,"tokens_out":26880,"duration_ms":260933,"significance":"The results are significant because they lower the smoothness threshold for Monte Carlo rates in L^p from B^1 to B^{1/2} and, more generally, show that deeper networks improve the rate for rough spectral Barron functions. The proofs are substantial: they combine multiscale cosine expansions, symmetrization with Khintchine inequalities, Rademacher complexity, and metric-entropy estimates, and the constants are explicit. The lower bound by oscillation counting has explicit constants and confirms sharpness. I found no error that touches the central claims; I also directly checked the composition identity in Lemma 3.2, and it holds, so the reader's flagged weakest assumption does not land. The main caveat is the paper's reliance on the companion paper [20] for Lemma 3.2 and Lemma 3.11, which should be verified as accepted and publicly available.","major_comments":[],"minor_comments":[{"comment":"The abstract states that 'the rates and prefactors in our estimates are dimension-free', but Theorem 3.6 and Corollary 3.9 have prefactors sqrt(1 + dL ln N) and (1 + dL ln N)^{sL}, respectively; please qualify the claim, for example by saying 'dimension-free up to logarithmic factors in d'.","section":"Abstract; Theorem 3.6; Corollary 3.9"},{"comment":"Corollaries 3.4 and 3.9, as well as the final unnumbered L^p improvement in Section 3.2, are stated without proofs ('details omitted for brevity'); since these corollaries are advertised results, please provide the full arguments or a detailed proof sketch with the relevant norm tracking.","section":"Section 3, Corollaries 3.4 and 3.9; final L^p improvement"},{"comment":"With |alpha_{xi,l,j}| <= 2^{1-l} pi and the prefactor 2^{(1+s)l}, the displayed constant 2^{1+sr} in the bound on ||F(.,xi,r)||_{Linfty(Omega)} appears to be missing a factor 2; the final rate is unaffected, but the estimate should be corrected.","section":"Proof of Theorem 2.4, bound on ||F(.,xi,r)||_{Linfty(Omega)}"},{"comment":"The optimization of \\tilde c and the conclusion 'Setting \\tilde c = 16.26^{-1} ... since c0 < 1' is too compressed; please include the intermediate inequalities that justify the final logarithmic factor, and fix the notation 16.26^{-1} to read 16.26^{-1}.","section":"Theorem 3.6, after Eq. (3.12)"},{"comment":"Lemma 3.2 and Lemma 3.11 are quoted from the companion paper [20]; because Lemma 3.2 is the key step that turns depth into the N^{-sL} rate, please either reproduce the proof or give a precise citation to the lemma's proof in [20].","section":"Lemma 3.2 and Lemma 3.11"},{"comment":"There are several typographical and formatting issues, including 'n-internals' in the proof of Theorem 3.10 (should be 'n intervals') and the rendering of 'H\\\"older' in a few places; these should be cleaned up before publication.","section":"General presentation"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript leans heavily on the authors' accepted paper [20] for two key lemmas (Lemma 3.2, the composition identity, and Lemma 3.11, the spectral estimate for Gaussian-windowed cosines). I verified the composition identity independently and it is correct; however, the editor should ensure that [20] is indeed in press and that its lemmas are available to readers, since the current paper's depth-dependence argument rests on that result. The abstract's 'dimension-free' claim should also be reconciled with the d-dependence in Theorem 3.6 before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. The paper does one real thing: it takes the authors' earlier L2 result for deep ReLU networks in spectral Barron space and pushes it to Lp for p >= 2 with a sqrt(p) prefactor, and to L-infinity with a sqrt(1 + dL ln N) prefactor, plus a matching Lp lower bound. That is a genuine extension, not a repackaging. The Lp proof via multiscale integral representation, symmetrization, and Khintchine is coherent; the uniform proof via Rademacher complexity and metric entropy is standard but carefully done. I checked the weakest point the reader flagged, Lemma 3.2 quoted from [20]: the composition identity gamma_{n2} composed with beta_{n1} equals gamma_{2n1n2} does hold, because beta_{n1} traverses each period twice and gamma is symmetric about 1/2. So that worry does not land. The lower bound adapts the oscillation-counting example from [9] to the Lp norm, and it gives the right N^{-sL} rate up to constants.\n\nThe paper leans heavily on the authors' own companion paper [20] for Lemma 3.2, Lemma 3.11, and the spectral Barron space machinery. That is fine as long as those results are correct; they appear to be, and they are cited properly. Self-citation is not a flaw here.\n\nThe soft spots are editorial, not technical. The abstract says the prefactors in the estimates are dimension-free, but Theorem 3.6 and Corollary 3.9 have the factor sqrt(1 + dL ln N), which depends on d and L. That overstates the uniform result. Several corollaries (3.4, 3.9, and the Lp improvement after Corollary 3.9) omit derivations with \"details are omitted for brevity.\" For a paper whose selling point is constants, this is frustrating; the omitted arguments are probably routine frequency splitting, but the authors should either spell them out or give precise pointers into [20].\n\nBottom line: the main theorems are sound, the new content is real, and the rate-sharpness picture is now clear for 0 < sL <= 1/2. I would send it to a serious referee. A revision should fix the abstract and fill in or carefully point to the corollary proofs. I would bring it to reading group, since it settles the Lp and uniform side of a question people in approximation theory care about.","headline":"Solid extension of the L2 theory to Lp and uniform rates for spectral Barron functions; the main theorems hold, and only the abstract wording and omitted corollary details need fixing.","tokens_in":24946,"tokens_out":1897,"would_cite":true,"duration_ms":20953,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["41A25","41A46","42A38","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that spectral Barron functions of smoothness $s$ are approximated by deep ReLU networks in $L^p$ at the sharp rate $N^{-sL}$ whenever $0<sL\\le 1/2$.","keywords":["spectral Barron space","deep neural networks","ReLU networks","Lp approximation","uniform approximation","sharp approximation rates","Monte Carlo rate","dimension-free constants"],"falsifier":"Evaluate the identity $\\gamma_{n_2}\\circ\\beta_{n_1}=\\gamma_{2n_1n_2}$ at the breakpoints $t=j/(2n_1n_2)$, $j=0,\\dots,2n_1n_2$, for small parameters such as $n_1=2$, $n_2=3$; a nonzero difference at any of those points would invalidate the width bound behind Theorem 3.1 and the claimed $N^{-sL}$ rate.","tokens_in":23883,"feed_emoji":"🧠","tokens_out":14448,"duration_ms":134946,"temperature":0.7,"pith_summary":"Approximation theory for neural networks traditionally required spectral Barron functions of smoothness at least $1$ to attain the Monte Carlo rate $N^{-1/2}$. This paper shows that smoothness $1/2$ is enough for shallow networks, and that with $L$ hidden layers any smoothness $s$ with $0<sL\\le 1/2$ yields the $L^p$ rate $N^{-sL}$ with a $\\sqrt{p}$ prefactor. The same order holds for uniform approximation, multiplied by at most $\\sqrt{1+dL\\ln N}$. A matching lower bound shows no ReLU or Heaviside network can do better than $N^{-sL}$ for worst-case functions in this range. Together these results fix the approximation power of deep ReLU networks on rough spectral Barron functions up to a logarithmic factor.","feed_headline":"Deep ReLU nets hit sharp rates for rough Barron functions","feed_subtitle":"Shallow nets hit the Monte Carlo rate at smoothness 1/2; depth extends it to rougher functions.","key_machinery":"The engine is the integral representation $f(x)=\\pi^2\\int_{\\mathbb{R}^d}\\int_0^1|\\widehat f(\\xi)|\\sin(\\pi r)\\gamma_{n_\\xi}(t_\\xi(x),r)\\,dr\\,d\\xi$, in which $\\gamma(t,r)$ is a ReLU-computable hat-shaped wavelet symmetric about $1/2$, and its repeated-period version $\\gamma_n(\\cdot,r)$ realizes $\\cos(2\\pi n t)$ through $\\pi^2\\int_0^1\\sin(\\pi r)\\gamma_n(t,r)\\,dr=\\cos(2\\pi n t)$. The composition identity $\\gamma_{n_2}\\circ\\beta_{n_1}=\\gamma_{2n_1n_2}$ (Lemma 3.2, from the companion paper [20]) lets each hidden layer multiply frequency, so choosing $n_i\\approx(1+|\\xi|_1)^{1/L}$ keeps the per-layer width at $4\\lceil(1+|\\xi|_1)^{1/L}\\rceil$ and converts a Monte Carlo sample of size $N$ into the global rate $N^{-sL}$. The $L^\\infty$ proof adds a Rademacher-complexity step with a metric-entropy bound on the class $x\\mapsto\\gamma_{n_\\xi}(t_\\xi(x),r)$, while the lower bound counts sign changes of $\\cos(2\\pi n x_1)$ to show that any $N$-unit-per-layer network must miss oscillation intervals.","core_discovery":"The central discovery is that the approximation rate for the spectral Barron space $\\mathscr{B}^s(\\Omega)$ is exactly $N^{-sL}$ in $L^p$ whenever the depth-smoothness product satisfies $0<sL\\le 1/2$, with constants independent of the dimension. Theorem 3.1 constructs an $(L,N)$-ReLU network with $\\|f-f_{L,N}\\|_{L^p(\\Omega)}\\le 13\\sqrt{p}\\|f\\|_{\\mathscr{B}^s(\\Omega)}N^{-sL}$ for $p\\ge2$; Theorem 3.6 gives the uniform analogue $45\\|f\\|_{\\mathscr{B}^s(\\Omega)}N^{-sL}\\sqrt{1+dL\\ln N}$; and Theorem 3.10 proves a matching lower bound of order $N^{-sL}$ for a worst-case function with $\\|f\\|_{\\mathscr{B}^s(\\Omega)}\\le 2+\\varepsilon$. Shallow sigmoidal networks already attain $N^{-s}$ for $0<s\\le1/2$, improving the earlier smoothness requirement $s\\ge1$ for the Monte Carlo rate and the earlier deep rate $N^{-sL/2}$ to the sharp exponent $sL$.","pith_inferences":["A testable extension of the same machinery: construct the paper's benchmark $f(x)=n^{-s}\\cos(2\\pi n x_1)e^{-\\pi|x|^2/R}$ and numerically measure the best $L^p$ error of $(L,N)$-ReLU networks; the predicted exponent $sL$ should appear for every $0<sL\\le1/2$, and any systematic slowing would point at the composition identity rather than the Monte Carlo step.","The paper's observation that sigmoidal activations reduce to Heaviside by shifting and scaling suggests the deep rates may transfer to tanh, softplus, ELU, and ReLU$^k$ networks; this transfer is not proved for depth here and is the natural next check.","Since the $L^\\infty$ upper bound carries a $\\sqrt{\\ln N}$ factor and the $L^p$ bounds do not, the open question left implicit is whether the sup-norm logarithmic factor is removable; that would need a matching $L^\\infty$ lower bound at logarithmically finer scales."],"forward_implications":["Functions in $\\mathscr{B}^s$ with $0<sL\\le1/2$, including functions too rough to be Hölder continuous of order above $s$, are approximated at order $N^{-sL}$ by deep ReLU networks with dimension-free prefactors.","Shallow networks with one hidden layer reach the Monte Carlo rate $N^{-1/2}$ in $L^p$ for $\\mathscr{B}^{1/2}$ functions, and the same rate in $L^\\infty$ up to a logarithmic factor.","The matching lower bound means the exponent $sL$ is optimal: no ReLU or Heaviside network with $L$ hidden layers and $N$ units per layer can approximate all of $\\mathscr{B}^s$ at a better worst-case order in this range.","Increasing depth directly improves the order for small smoothness: with $s=0.05$, the paper notes that 2, 6, and 10 hidden layers give rates approaching $N^{-1/10}$, $N^{-3/10}$, and $N^{-1/2}$, respectively.","The uniform bound loses only a $\\sqrt{dL\\ln N}$ factor, so the dimension-free character of the Monte Carlo rate survives in the sup norm."],"supporting_citations":[{"why":"Companion theory: supplies the spectral Barron norm and embedding, the composition identity $\\gamma_{n_2}\\circ\\beta_{n_1}=\\gamma_{2n_1n_2}$, the $L^2$ $N^{-sL}$ bound, and the spectral-norm lemma for the lower-bound example.","marker":"[20]"},{"why":"Supplies the deep integral representation and Lemma 3.3 expressing $\\cos(2\\pi nt)$ through $\\gamma_n$, whose earlier $L^2$ rate $N^{-sL/2}$ is improved here.","marker":"[9]"},{"why":"Provides the multiscale expansion and the first $\\mathscr{B}^{1/2}$ $L^2$ $N^{-1/2}$ shallow result that the $L^p$ shallow theorem refines with dimension-free constants.","marker":"[36]"},{"why":"Gives the shallow sigmoidal $L^p$ approximation for $\\mathscr{B}^1$ functions used on the low-frequency part in the corollaries.","marker":"[24]"},{"why":"Provides the $L^\\infty$ approximation bound for $\\mathscr{B}^1$ ReLU networks used for the low-frequency component of the uniform estimate.","marker":"[10]"},{"why":"Gives the optimal Khintchine constants; Lemma 2.8 sharpens them into the explicit $\\sqrt{p}$ prefactor.","marker":"[16]"},{"why":"Supplies the metric-entropy bound on Rademacher complexity (Lemma 3.8) that converts the $L^2$ Monte Carlo estimate into the $L^\\infty$ bound.","marker":"[31]"},{"why":"Provides the sign-change counting lemma used in the lower bound to show any $(L,N)$-network misses oscillation intervals.","marker":"[39]"},{"why":"Cube-slicing lemma used in Lemma 2.2 to control the $L^p$ distance between a sigmoidal unit and its Heaviside limit.","marker":"[4]"}],"fun_headline_variants":["Half smoothness, same shallow-net rate","Depth-smoothness product sets sharp net rate","Smoothness 1/2 suffices for MC net rate","N^{-sL} tight for rough Barron nets","Depth extends MC rate to rougher Barron nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything in the deep part rests on an exact frequency-doubling identity for the paper's repeated wave shapes: composing a wave with a triangle wave must reproduce the wave with twice as many periods on the whole interval $[0,1]$, including at the breakpoints; if that identity fails at a single point, the per-layer width bound and the $N^{-sL}$ rate no longer follow.","fun_headline_variants_meta":{"raw":{"variants":["Half smoothness, same shallow-net rate","Depth-smoothness product sets sharp net rate","Smoothness 1/2 suffices for MC net rate","N^{-sL} tight for rough Barron nets","Depth extends MC rate to rougher Barron nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000696,"raw_usage":{"total_tokens":3186,"prompt_tokens":1023,"completion_tokens":2163,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":2085}},"tokens_in":639,"tokens_out":2163,"duration_ms":22614,"temperature":1.0,"reasoning_tokens":2085,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:02:26.679776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the identity $\\gamma_{n_2}\\circ\\beta_{n_1}=\\gamma_{2n_1n_2}$ at the breakpoints $t=j/(2n_1n_2)$, $j=0,\\dots,2n_1n_2$, for small parameters such as $n_1=2$, $n_2=3$; a nonzero difference at any of those points would invalidate the width bound behind Theorem 3.1 and the claimed $N^{-sL}$ rate.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the deep integral representation and Lemma 3.3 expressing $\\cos(2\\pi nt)$ through $\\gamma_n$, whose earlier $L^2$ rate $N^{-sL/2}$ is improved here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the multiscale expansion and the first $\\mathscr{B}^{1/2}$ $L^2$ $N^{-1/2}$ shallow result that the $L^p$ shallow theorem refines with dimension-free constants."},{"cited_title":"Makovoz, Random approximants and neural networks , J","cited_arxiv_id":null,"evidence_quote":"Gives the shallow sigmoidal $L^p$ approximation for $\\mathscr{B}^1$ functions used on the low-frequency part in the corollaries."},{"cited_title":"Caragea, P","cited_arxiv_id":null,"evidence_quote":"Provides the $L^\\infty$ approximation bound for $\\mathscr{B}^1$ ReLU networks used for the low-frequency component of the uniform estimate."},{"cited_title":"Haagerup, The best constants in the Khintchine inequality , Studia Math","cited_arxiv_id":null,"evidence_quote":"Gives the optimal Khintchine constants; Lemma 2.8 sharpens them into the explicit $\\sqrt{p}$ prefactor."},{"cited_title":"Rebeschini, Lecture notes in algorithmic foundations of learning: Cove ring numbers bounds for Rademacher complexity","cited_arxiv_id":null,"evidence_quote":"Supplies the metric-entropy bound on Rademacher complexity (Lemma 3.8) that converts the $L^2$ Monte Carlo estimate into the $L^\\infty$ bound."},{"cited_title":"Telgarsky, Beneﬁts of depth in neural networks , Proceedings of Machine Learning Re- search 49 (2016), 1517–1539","cited_arxiv_id":null,"evidence_quote":"Provides the sign-change counting lemma used in the lower bound to show any $(L,N)$-network misses oscillation intervals."},{"cited_title":"Ball, Cube slicing in Rn, Proc","cited_arxiv_id":null,"evidence_quote":"Cube-slicing lemma used in Lemma 2.2 to control the $L^p$ distance between a sigmoidal unit and its Heaviside limit."}],"review_version":1}