{"id":"9a0d5a07-2dc5-4119-a982-53eca5c48ddc","arxiv_id":"2605.31152","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Deep ReLU networks approximate anisotropic Besov functions at rate O((WL)^(-2\\tilde{s})) and mixed-smooth Besov functions at rate O((WL)^(-2s)) up to logs, with matching lower bounds up to logs.","lead":"This paper proves new approximation and learning rates for deep ReLU networks on functions with different degrees of smoothness in different directions, and on functions with mixed smoothness. The rates can avoid the usual exponential slowdown with dimension when the right anisotropic or mixed-smooth structure is assumed.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the central super-approximation proofs for anisotropic and mixed Besov spaces hold up under scrutiny.","rationale":"The paper's central approximation results are supported by detailed, internally consistent proofs. The super rate (WL)^{-2\\tilde{s}} for anisotropic Besov and (WL)^{-2s} (up to logs) for mixed Besov follow from a careful adaptive multiscale decomposition combined with the interpolation capacity of fully-connected ReLU networks. I checked the load-bearing steps: the local Whitney estimate, the conversion to averaged moduli, the summing property, the coefficient norm bounds, and the final width/depth trade-off. The averaged-modulus equivalence flagged by the reader is a classical coordinate-wise fact; the proof in the manuscript is abbreviated but the constants are uniform and the condition is satisfied. The lower-bound arguments using pseudo-dimension and metric entropy are standard and correctly applied. The minor gaps—an unproved external interpolation lemma and a 'cumbersome calculation' in the composition-model lower bound—do not undermine the main theorems and are addressable. Therefore I do not identify a load-bearing concern; the verdict remains CONDITIONAL as given by the reader.","tokens_in":45245,"tokens_out":45406,"duration_ms":391675,"concrete_test":"Run an independent numerical check of Lemma 5.9 for d=2, k=2, q=2 on random anisotropic functions: compute ω and ω̃ on a sequence of rectangles with side lengths shrinking geometrically, and verify the ratio ω/ω̃ remains bounded as t approaches δ_j/(4k). If the ratio diverges for some configuration, the decomposition in Lemma 5.10 would need a different modulus; otherwise the concern is closed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After a detailed pass, I find no load-bearing flaw in the central claims (Theorems 3.1 and 3.3). The reader's weakest assumption concerned Lemma 5.9's averaged-modulus equivalence and its use in Lemma 5.10. This concern does not land: the ordinary modulus ω and the averaged modulus ω̃ are equivalent coordinate-wise with constants depending only on k, not on the rectangle side lengths; the proof's condition mismatch (4k vs 4k^2) is a typo, and the application uses t = λ b^{-ℓ_j} with λ = 1/(5k_j), which satisfies t ≤ δ_j/(4k_j). The summing property of ω̃ is valid via Minkowski's integral inequality: the ℓ_q norm of the vector of local averaged moduli is bounded by the global averaged modulus. The rest of the decomposition (Lemma 5.10), the network interpolation (Lemma 5.7), and the width/depth balancing (Propositions 5.11 and 5.13) are internally coherent. The remaining issues—reliance on external Lemma 5.6 and the unshown 'cumbersome calculation' in Theorem 3.6—are secondary and do not affect the main approximation theorems.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies approximation and learning of anisotropic and mixed-smooth Besov functions by fully-connected deep ReLU networks. The central upper bounds are Theorem 3.1, which gives an Lp approximation rate of order (WL)^{-2\\tilde{s}} for anisotropic Besov spaces under \\tilde{s} > 1/q - 1/p, and Theorem 3.3, which gives the rate (WL)^{-2s}(\\log W \\log L)^{(d-1)(2s+1)} for mixed-smooth Besov spaces under s > 1/q - 1/p. Theorem 3.5 extends the construction to compositions of anisotropic Besov functions. Theorems 3.2, 3.4 and 3.6 provide pseudo-dimension/entropy-based lower bounds showing that these rates are optimal up to logarithmic factors. Section 4 applies the approximation bounds to least-squares regression and derives minimax-optimal learning rates, up to logarithms, for the corresponding smoothness classes. The proofs are built on anisotropic multiresolution partitions, averaged moduli of smoothness, piecewise-polynomial decompositions, and the interpolation networks from Yang (2025).","tokens_in":45520,"tokens_out":36810,"duration_ms":337640,"significance":"If the results stand, they are a substantial advance: they show that the good approximation rates of deep ReLU networks for anisotropic and mixed smoothness do not require sparse architectures, and they provide explicit width/depth trade-offs. The extension from isotropic to anisotropic approximations is nontrivial, and the use of averaged moduli to sum local Whitney estimates is well suited to the problem. The paper is also careful about constants and about the distinction between the parallel and sequential summation of networks. The lower-bound arguments are standard but are adapted cleanly to the anisotropic and mixed cases. Overall, the central claims of Theorems 3.1 and 3.3 appear sound; the issues I found are local and fixable.","major_comments":[],"minor_comments":[{"comment":"The statement requires 0 < t ≤ δ_j/(4k), but the proof only establishes the equivalence for 0 < t ≤ δ_j/(4k^2) (see the constraint 'h ≤ t/2 ≤ δ_j/(8k^2)' and the final line 't ≤ δ_j/(4k^2)'). In Lemma 5.10 the argument applies the lemma with t = λ b^{-ℓ_j}, λ = min_j 1/(5k_j), which satisfies the stated 4k condition but not the proof's 4k^2 condition when k_j > 1. This mismatch should be fixed explicitly: either complete the proof for the stated 4k condition or, alternatively, replace λ by min_j 1/(5k_j^2) in Lemma 5.10; the rest of the estimates are unaffected since the new constants still depend only on p,q,s,d,b.","section":"Lemma 5.9"},{"comment":"The sentence 'through a cumbersome calculation, one can verify that x^{eγ_m} ∈ B^{s_m}_{q_m}([0,1])' is an omitted proof of a claim used to construct the lower-bound family. The endpoint case eγ_m = s_m - 1/q_m with r = ∞ is plausible via the critical embedding, but the calculation should be written out, especially since Theorem 3.6 supports the claimed near-optimality of Theorem 3.5.","section":"Theorem 3.6 proof"},{"comment":"The application of Corollary 5.2 to the N_{ℓ*,d} networks G_i is terse. A single application of the corollary gives only one of the two polynomial factors: width O(N b^α) with depth O(b^β), or width O(b^α) with depth O(N b^β). To obtain simultaneously W ≤ C α^{d-1} b^α and L ≤ C β^{d-1} b^β, the networks must be grouped into a two-dimensional grid (N_1 × N_2 with N_1 N_2 ≥ N_{ℓ*,d}). Please spell out this grouping.","section":"Proposition 5.13"},{"comment":"Several typos should be corrected: 'supper approximation' → 'super approximation'; 'extent this result' → 'extend this result'; 'there exits f' → 'there exists f'. In Section 3.2, after Theorem 3.4, 'the upper bound in Theorem 3.1' should be 'Theorem 3.3'. In the proof of Theorem 3.5, g_m should map to R^{d_{m+1}} rather than R^{d_m}, and the width bound should accordingly use max_{2≤m≤M} d_{m+1} W_0.","section":"Typos and notation"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new result is that fully-connected ReLU networks achieve (WL)^{-2\\tilde{s}} for anisotropic Besov spaces and (WL)^{-2s} up to logs for mixed Besov spaces, without sparsity constraints. That extends Suzuki and Nitanda's sparse-network rates and Yang's isotropic fully-connected result, and it matters because standard architectures get the same dimension-robust behavior without needing to know where nonzero parameters are. The proofs are detailed and internally consistent: Lemma 5.7 is the key workhorse, and the width/depth balancing in Propositions 5.11 and 5.13 is careful. I checked the averaged-modulus concern from your notes: the constants in Lemma 5.9 depend only on k, not on rectangle side lengths, and the 4k vs 4k^2 mismatch is a typo. So the central approximation claims look solid. The soft spots are secondary. Theorem 3.6, the lower bound for the deep composition model, has an unshown 'cumbersome calculation' and leans on external pseudo-dimension/entropy estimates. That is a real presentation gap, but it does not affect Theorems 3.1 and 3.3. The abstract says minimax optimal rates 'for a wide range of smooth function classes,' which is too broad: for the composition model the upper bound uses s* and the lower bound uses s**, and they match only when all q_m = infinity. The authors acknowledge this in the text, so it is an overstatement in the abstract rather than a fatal flaw. The learning section is routine application of known oracle inequalities. The citation pattern is fine - Yang 2025 and Suzuki/Nitanda are cited appropriately, and the claimed novelty is real. Who is this for? Approximation theorists and people working on neural network expressivity. It deserves a serious referee; the main results are likely correct and are a within-subfield advance. I would send it to review, and I'd ask the authors to fill in the omitted calculation and to soften the abstract's minimax claim.","headline":"Fully-connected ReLU nets get super rates for anisotropic and mixed Besov spaces, and the main proofs hold up; the abstract overclaims minimaxity for the composition model and one lower bound has an unshown calculation.","tokens_in":637,"tokens_out":1410,"would_cite":true,"duration_ms":28046,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["41A25","41A46","46E35","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep ReLU networks approximate anisotropic and mixed-smooth functions at rates set by the harmonic mean of the coordinate smoothnesses.","keywords":["deep ReLU networks","anisotropic Besov spaces","mixed smoothness","approximation rates","curse of dimensionality","minimax optimality","nonparametric regression","piecewise polynomial approximation"],"falsifier":"Numerically compute the ratio of the averaged modulus $\\tilde{\\omega}(f,t)$ to the ordinary modulus $\\omega(f,t)$ on a sequence of rectangles with increasingly extreme side ratios, with $t$ just below $\\delta_j/(4k^2)$, for a function such as $f(x)=x_1^{s_1}x_2^{s_2}$ with widely separated $s_1,s_2$. If the ratio grows without bound as the side ratio grows, Lemma 5.9 fails and the rate is not supported. Alternatively, compute the coefficient bound $||a_\\ell||_q$ for such an $f$ and check whether it really follows $b^{\\ell/q - \\tilde{s}\\ell}$ or picks up a factor growing with $s_1/s_2$.","tokens_in":45130,"feed_emoji":"🧠","tokens_out":4801,"duration_ms":46096,"temperature":0.7,"texified_at":"2026-08-05T21:07:25.025808+00:00","pith_summary":"This paper asks how accurately deep ReLU networks can approximate smooth functions when different coordinate directions have different smoothness, or when smoothness is mixed across coordinates. It claims that for anisotropic Besov spaces the approximation error is bounded by $C (WL)^{-2\\tilde{s}}$ where $\\tilde{s}$ is the harmonic mean of the per-coordinate smoothness exponents, and that for mixed-smooth Besov spaces the same rate holds up to logarithmic factors. Because $\\tilde{s}$ can be independent of the ambient dimension, these rates escape the curse of dimensionality whenever the function is sufficiently smooth in most directions. The paper also shows the rates are near-optimal and that fully-connected networks reach minimax statistical rates for nonparametric regression on these classes. A sympathetic reader should take the paper as establishing that adaptive, anisotropic piecewise-polynomial decompositions can be realized by reasonably sized fully-connected ReLU networks.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":6642,"prompt_tokens":895,"completion_tokens":5747,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":895,"completion_tokens_details":{"reasoning_tokens":4913}},"feed_headline":"Harmonic mean of directional smoothness sets ReLU approximation rate","feed_subtitle":"Anisotropic Besov classes are approximated at (WL)^-2s-tilde; mixed-smooth classes match, up to log factors.","key_machinery":"The proof rests on an adaptive multiscale anisotropic grid: at scale $\\ell$, each coordinate direction j is partitioned into about $b^{\\ell \\tilde{s}/s_j}$ intervals, so the number of rectangles grows like $b^\\ell$ regardless of dimension. The averaged modulus of smoothness converts local Whitney approximation errors on these rectangles into a global coefficient bound $||a_\\ell||_q \\lesssim b^{(1/q-\\tilde{s})\\ell}$, which controls the sparse coefficients fed to a network interpolation lemma for discretized data. A base-prime trick constructs several approximants on slightly shifted grids and uses an order-statistic network to select the best one, removing the 'trifling' boundary region. The width/de","core_discovery":"The central claim is that the 'super approximation rate' known for isotropic Besov functions extends to anisotropic and mixed smoothness without sparse architectures. For anisotropic Besov space $B^s_q([0,1]^d)$ with mean smoothness $\\tilde{s} = (\\sum_j s_j^{-1})^{-1}$, if $\\tilde{s} > 1/q - 1/p$, the network class $NN(W,L)$ satisfies $\\sup_{f: ||f||\\le 1} \\inf_g ||f-g||_{L^p} \\le C (WL)^{-2\\tilde{s}}$. For mixed-smooth Besov space $MB^s_q$, if $s > 1/q - 1/p$, the rate is $(WL)^{-2s} (\\log W \\log L)^{(d-1)(2s+1)}$. These rates are optimal up to logarithmic factors, and the same machinery yields rates for compositions of anisotropic Besov functions and for least-squares learning.","pith_inferences":["If the constants in the key coefficient estimate turn out to depend badly on the dimension d or on the spread of the s_j, the practical benefit in high dimensions would shrink even though the formal rate is dimension-free; the paper leaves this dependence unquantified.","The same anisotropic-grid averaging argument may extend to smoothness classes with mixed coordinate interactions beyond product Besov spaces, such as Triebel–Lizorkin spaces with dominating mixed smoothness.","A direct testable corollary: for additive models or tensor-product functions listed in Remark 2.1, the network approximation error should behave as (WL)^{-2s} with logarithmic factors; this can be checked numerically with small networks.","The technique suggests a recipe for other univariate nonlinear approximators: any basis that yields Whitney estimates and averaged-modulus sum inequalities can be converted into a deep ReLU approximation bound."],"forward_implications":["Anisotropic Besov approximation at rate (WL)^{-2\\tilde{s}} is optimal up to a logarithmic factor; with bounded width the rate becomes L^{-2\\tilde{s}}.","Mixed-smooth Besov approximation reaches (WL)^{-2s} up to logs, improving on sparse-network results and matching lower bounds up to logarithmic factors.","Fully-connected ReLU networks — no sparsity constraints — attain minimax regression rates n^{-2\\tilde{s}/(2\\tilde{s}+1)} (and the mixed-smooth analogue) up to logs.","Compositions of anisotropic Besov functions are approximated at a rate determined by the smoothness index s^* = \\min_m \\tilde{s}_m \\prod_{k>m} \\gamma_k, near-optimal when all q_m = \\infty.","Existing isotropic Sobolev, Hölder, and Besov bounds appear as the special case s = (s_0,\\ldots,s_0), giving \\tilde{s} = s_0/d."],"fun_headline_variants":["Deep ReLU nets match minimax rates for anisotropic and mixed smoothness","Harmonic mean of directional smoothness sets optimal ReLU rate","Anisotropic Besov spaces approximated at optimal ReLU rates","Mixed smoothness now within reach for deep ReLU nets","ReLU networks achieve optimal rates beyond isotropic smoothness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole argument leans on a single estimate asserting that an averaged smoothness measure stays comparable to the usual modulus of smoothness with a constant that does not blow up as rectangles become long and thin; if that constant actually depends on the coordinate smoothness parameters, the coefficient bounds and the final $(WL)^{-2\\tilde{s}}$ rate collapse.","fun_headline_variants_meta":{"raw":{"variants":["Deep ReLU nets match minimax rates for anisotropic and mixed smoothness","Harmonic mean of directional smoothness sets optimal ReLU rate","Anisotropic Besov spaces approximated at optimal ReLU rates","Mixed smoothness now within reach for deep ReLU nets","ReLU networks achieve optimal rates beyond isotropic smoothness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001488,"raw_usage":{"total_tokens":5893,"prompt_tokens":909,"completion_tokens":4984,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":4899}},"tokens_in":653,"tokens_out":4984,"duration_ms":30720,"temperature":1.0,"reasoning_tokens":4899,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T12:45:56.874721+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Numerically compute the ratio of the averaged modulus $\\tilde{\\omega}(f,t)$ to the ordinary modulus $\\omega(f,t)$ on a sequence of rectangles with increasingly extreme side ratios, with $t$ just below $\\delta_j/(4k^2)$, for a function such as $f(x)=x_1^{s_1}x_2^{s_2}$ with widely separated $s_1,s_2$. If the ratio grows without bound as the side ratio grows, Lemma 5.9 fails and the rate is not supported. Alternatively, compute the coefficient bound $||a_\\ell||_q$ for such an $f$ and check whether it really follows $b^{\\ell/q - \\tilde{s}\\ell}$ or picks up a factor growing with $s_1/s_2$.","supporting_citations":[],"review_version":2}