{"id":"f9f83ce0-dc09-42a0-8376-974e52870e30","arxiv_id":"2608.06687","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Linearized ReLU^k discrete least squares on the sphere achieves the optimal rate n^{-(r-s)/d} with m roughly n deterministic samples under a parity condition and quasi-uniform points.","lead":"The paper proves that deterministic collocation with only about as many sample points as network parameters reaches the optimal approximation rate for linearized ReLU^k least-squares solvers of elliptic equations on spheres. This gives a rigorous theoretical basis for a common numerical strategy that previously lacked sample guarantees.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4's deterministic rate rests on an exact low-degree coefficient identity imported from [49] without proof; the identity is used in (4.23)-(4.25), though a modified bound may patch the proof.","rationale":"The paper is carefully written and honest about limitations; the spherical proof is mostly self-contained except for Theorem 6. The central claim, that deterministic collocation with m roughly n attains the continuous rate, is plausible and supported by a coherent argument once the comparison network is available. The reader's weakest assumption is exactly the unproven low-degree identity, and I agree this is the softest part. My analysis also shows the identity is convenient rather than indispensable: the extra polynomial mismatch term can be absorbed at the same order. Because the fix is local and the external result is likely available in the cited prior work, the appropriate verdict remains conditional, not reject or accept. The empty-range typo in (4.26) and the missing final step from Theorem 3 to Theorem 1 are expositional and easily repaired.","tokens_in":32835,"tokens_out":16491,"duration_ms":154953,"concrete_test":"Obtain [49] and check whether equations (4.12)-(4.13) there construct u_n satisfying both (4.5) and the exact identity (4.6) for all nu <= C_1 h^{-1}, with constants independent of n and Theta_n. Independently, re-run the proof of Theorem 4 without invoking P_q(f-L_beta u_n)=0, instead bounding ||P_q L_beta(u_n-u)||_{L^2} by (4.16); if the final h^{-s} h^r estimate still closes, the identity is not load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest link is Theorem 6, quoted from the authors' prior paper [49] and stated in a stronger form than the cited theorem. The text says the comparison network u_n satisfies both the approximation estimate (4.5) and the exact low-degree identity (4.6), with the identity following from [49, (4.12)-(4.13)] but not proved here. The proof of Theorem 4 uses (4.6) to conclude P_q u_n = P_q u, hence P_q(f-L_beta u_n)=0, in equations (4.23) and (4.25). If (4.6) fails at the stated strength, that equality is unavailable. This does not immediately refute the theorem: since P_q is an L^2 contraction, ||P_q L_beta(u_n-u)||_{L^2} <= ||L_beta(u_n-u)||_{L^2} ~ h^r ||f||_{H^r} by (4.16), so the argument can be patched by carrying this extra term. The concern is therefore a gap in the presented derivation rather than a counterexample to the rate, but the strongest claim of the paper depends on an unverified external result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops deterministic and randomized discrete residual least-squares theory for elliptic spectral equations L_beta u = f on the unit sphere, using linearized ReLU^k network spaces with fixed, antipodally quasi-uniform parameter sets. The main deterministic result (Theorem 4) states that for any quasi-uniform collocation set with m >= C h^{-d} (so m proportionally n in the quasi-uniform case), the empirical residual minimizer achieves the continuous approximation rate h^{r-s}, equivalently n^{-(r-s)/d}. Theorem 5 gives a high-probability analogue for i.i.d. uniform collocation points, with a log factor and an arbitrarily small smoothness loss. The central technical novelty is a Bernstein inequality for ReLU^k network spaces (Theorem 1), with an explicit dependence on the antipodal separation distance, together with a residual-space version for L_beta L_n^k. Section 5 provides an affine-network interpretation for the order-zero case on bounded Lipschitz domains, and Section 6 reports numerical experiments consistent with the predicted rates.","tokens_in":33067,"tokens_out":21613,"duration_ms":197418,"significance":"If the main results are valid, the deterministic sample complexity m ~ n is a strong and interesting statement: it shows that equal-weight collocation with only a constant oversampling factor can match the optimal continuous Sobolev approximation rates for linearized ReLU^k networks, without cubature weights or polynomial exactness. The Bernstein inequality with explicit geometric constants is a useful independent contribution, and its inverse-approximation corollary is a meaningful addition to the literature. The paper is also commendably careful about the scope of the bounded-domain extension, explicitly stating that it does not prove a direct collocation theorem on general domains. The main caveat is that the deterministic proof relies on a strengthened comparison theorem imported from the authors' prior work, which is not proved in this manuscript; this limits the certainty of the headline claim until the gap is closed.","major_comments":[{"comment":"The deterministic rate in Theorem 4 depends on the exact low-degree identity (4.6), namely \\hat u_n(ν,ℓ)=\\hat u(ν,ℓ) for ν≤C_1 h^{-1}. The paper states that the cited theorem in [49] is weaker, and that the identity follows from [49, (4.12)–(4.13)], but the identity is not proved here. This is load-bearing: it is used to conclude P_q f = P_q L_β u_n, which underlies the decomposition in (4.23) and the equality in (4.25). As written, Theorem 4 is conditional on an unverified strengthening of an external result. I do not regard this as a counterexample, because the argument can likely be repaired by estimating \\|P_q^c(L_β u_n - f)\\|_{L^2} ≤ \\|L_β u_n - f\\|_{L^2} ≲ h^r \\|f\\|_{H^r} and carrying the extra term through (4.24)–(4.25), but the proof should present this repair or prove (4.6).","section":""},{"comment":"The Bernstein inequality, which is the key analytical ingredient, relies on the asymptotic estimate (3.18) for ξ_α^{(ν)}(t), quoted from [49, Lemma 3.1] without proof. Since this estimate controls the localization estimate in Lemma 1 and hence the diagonal-dominance argument in Lemma 2, the main inequality is not fully self-contained. I suggest including a short derivation of (3.18) from gamma-function asymptotics, or at least reproducing the cited lemma in full, so that the novelty of Theorem 1 can be verified independently of the authors' prior preprint.","section":""}],"minor_comments":[{"comment":"The first branch in both displays reads \"d/2 < r ≤ d/2\", which is an empty range; it should read \"d/p < r ≤ d/2\" as in the statement of Theorem 4 and the abstract.","section":""},{"comment":"The condition \"d/p < r ≤ d/2, p > 2\" should be read as \"there exists p > 2 with d/p < r\"; the quantifier over p should be made explicit in the theorem statement and in the abstract.","section":""},{"comment":"The column heading \"reference rate\" for the cube experiments lists the spherical rates from Theorem 4 even though the text states they are not theoretical rates for Q^d; relabel the column as \"spherical reference\" to avoid confusion.","section":""},{"comment":"The quasi-uniform point set in Theorem 7 is a set of spherical collocation points, not samples in Ω; please state this explicitly in the theorem hypothesis, since the surrounding text is clear but the theorem statement alone could be misread.","section":""}],"recommendation":"major_revision","confidential_remarks":"The paper's headline deterministic result depends on [49], an arXiv preprint by the same authors, in two places: the comparison theorem (Theorem 6) and the estimate (3.18). The former is explicitly stated in a stronger form than the cited theorem. I recommend asking the authors to either supply a proof of the strengthening or revise the proof of Theorem 4 so that it relies only on the approximation estimate (4.5). This is fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. The paper delivers a real result: deterministic residual least squares for ReLU^k networks on the sphere with m ~ n samples and optimal rates, powered by a Bernstein inequality for the linearized space with explicit dependence on antipodal separation. The Bernstein inequality is the genuine new ingredient; the sample complexity theorem is its payoff. The paper is also honest about its limits: the bounded-domain construction is indirect and confined to the order-zero case, and the numerics are called finite-range observations.\n\nWhat it does well beyond the headline: the inverse approximation theorem follows naturally, the proof of the Bernstein inequality via localized Jacobi kernels and diagonal dominance is detailed and coherent, the random sampling result is a natural extension, and the experiments match the predicted exponents with reported diagnostics.\n\nThe main soft spot is one load-bearing dependency. Theorem 4 uses a comparison network from the authors' prior paper [49] that is stated in stronger form than the cited theorem: it adds exact low-degree coefficient matching, used to get P_q(f - L_beta u_n) = 0. That identity is not proved in this paper. If [49] does not actually establish it, the proof as written has a gap. The stress-test note is right that the rate can be patched—P_q is an L^2 contraction, so the un-cancelled term is absorbed by the existing approximation error—but the manuscript should either prove the identity or give the precise location in [49]. There is also a typo in (4.26)–(4.27): the condition reads 'd/2 < r <= d/2', which is empty; it should be 'd/p < r <= d/2'. The absence of code is minor given the experimental detail.\n\nOverall the math is coherent and the central claim is solid. This paper is for numerical analysts working on neural-network PDE solvers and approximation theorists interested in network Bernstein inequalities. It deserves a serious referee; after the dependency and typo are fixed, it should be accepted.","headline":"A genuinely new Bernstein inequality and a deterministic m~n sample bound for ReLU^k residual least squares on the sphere; the main caveat is an unproved stronger comparison theorem imported from the authors' prior work.","tokens_in":33572,"tokens_out":5635,"would_cite":true,"duration_ms":50701,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["41A25","41A30","65D15","65N12","65N35"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that empirical least squares over deterministic collocation points achieves the optimal $n^{-(r-s)/d}$ approximation rate for linearized ReLU$^k$ networks on the sphere, with a sample size proportional to the number of…","keywords":["neural network approximation","empirical least squares","deterministic collocation","Bernstein inequality","ReLU^k activation","spherical harmonics","Sobolev spaces","elliptic spectral equations"],"falsifier":"Inspect the construction behind the comparison-network theorem: compute the low-degree spherical harmonic coefficients of the comparison network and compare them with the target's coefficients for all degrees up to $C_1 h^{-1}$. If the identity $\\widehat{u_n}(\\nu,\\ell)=\\widehat{u}(\\nu,\\ell)$ fails at any degree in that range, then the proof's step $P_q L_\\beta u_n = P_q f$ no longer holds and the claimed rate is unsupported. An independent check would run the discrete least-squares problem with $m$ proportional to $n$ on a smooth parity-compatible right-hand side and test whether the empirical error decays at the predicted $n^{-(r-s)/d}$ rate over a wide range of $n$.","tokens_in":32626,"feed_emoji":"🌐","tokens_out":11363,"duration_ms":89595,"temperature":0.7,"pith_summary":"The paper proves that for elliptic equations $L_\\beta u = f$ on the sphere solved by linearized ReLU$^k$ networks, empirical least squares over deterministic collocation points reaches the same optimal approximation rate as the continuous formulation, provided the number of sampling points $m$ is proportional to the number of network parameters $n$ and the target has the parity compatible with the activation. The trial space consists of ridge functions $\\sigma_k(\\theta_j \\cdot \\eta)$ with parameters on the sphere, and $L_\\beta$ is a positive elliptic spectral multiplier of order $\\beta$. The key stability tool is a Bernstein inequality controlling high-order Sobolev norms of network functions by lower-order norms through the antipodal separation distance of the parameters. Because $m\\asymp n$ suffices, the method bypasses heavy oversampling and quadrature. The paper also gives a near-optimal high-probability bound for i.i.d. uniform samples and, for $\\beta=0$, an affine-network version on bounded Lipschitz domains through a constructive lifting.","feed_headline":"Optimal rates need only m ≈ n deterministic samples","feed_subtitle":"For elliptic ReLU^k network equations on the sphere, discrete least squares matches continuous approximation order.","key_machinery":"The load-bearing object is the Bernstein inequality for linearized ReLU$^k$ spaces on the sphere: for $0\\le s<r<k+1/2$, every $v_n\\in L_n^k$ satisfies $\\|v_n\\|_{H^r} \\lesssim \\underline{h}^{-(r-s)}\\|v_n\\|_{H^s}$, where $\\underline{h}$ is the antipodal separation distance of the parameters. Its proof fills the spectral gaps of $\\sigma_k$ with a surrogate kernel $\\phi_k$, represents Sobolev seminorms as quadratic forms built from localized dyadic kernels, and establishes diagonal dominance of the resulting matrices under antipodal separation. Applied to the residual space $L_\\beta L_n^k$, the same inequality gives the inverse estimate that turns empirical residual control into Sobolev error control. A companion comparison-network theorem supplies an approximant whose low-degree spherical harmonic coefficients match the target exactly, which is what lets the proof identify low-frequency parts of the residual before applying the Bernstein estimate.","core_discovery":"Theorem 4 is the central claim: for $k>\\beta+(d-1)/2$, $r\\le (d+2k+1)/2-\\beta$, and $0\\le s<\\min\\{r,k+1/2-\\beta\\}$, if the network parameter set is antipodally quasi-uniform and the collocation set is quasi-uniform with $m\\ge C_2 h^{-d}$, then the empirical residual minimizer satisfies $\\|u-u_{n,m}\\|_{H^{s+\\beta}} \\eqsim \\|f-L_\\beta u_{n,m}\\|_{H^s}$, bounded above by $h^{r-s}$ times a Sobolev norm of $f$. With $m\\asymp n$ this becomes the optimal rate $n^{-(r-s)/d}$ in the target norm. The paper's point is that the unregularized linear least-squares problem with deterministic points retains the full approximation power of the continuous problem. In the order-zero case $\\beta=0$ the same estimate recovers ordinary discrete least-squares approximation, and the authors make clear that the bounded-domain version is an indirect lifting-based consequence, not a direct Euclidean-domain collocation theorem.","pith_inferences":["The paper explicitly restricts the bounded-domain construction to the order-zero case and does not claim a direct collocation theory on Euclidean domains; extending the Bernstein argument to affine networks with boundary residuals is a natural next step, but the paper leaves it open.","If the comparison-network identity can be established beyond antipodally quasi-uniform parameter sets, the same proof would carry the optimal-rate conclusion to other geometries; the paper does not assert this.","The explicit dependence of the sample threshold on $h^{-d}$ suggests that the optimal rate is tied to parameter geometry: poorly separated parameters force more collocation points, so the practical claim is about well-conditioned parameter sets.","The reported experiments cover finite refinement ranges; a sharper test of the theory would measure slopes at considerably larger $n$ to separate asymptotic rates from pre-asymptotic transients."],"forward_implications":["For any positive elliptic spectral multiplier $L_\\beta$ on the sphere, residual least squares with $m\\asymp n$ deterministic quasi-uniform collocation points achieves the optimal rate $n^{-(r-s)/d}$ for parity-compatible data of Sobolev smoothness $r$.","The Bernstein inequality gives an inverse approximation theorem: if a function is approximated at rate $n^{-(r-s)/d}$ in $H^s$ by these networks, then it belongs to $H^\\alpha$ for every $\\alpha<r$.","For $\\beta=0$, the theory covers ordinary discrete least-squares approximation, and the lifting construction yields affine ReLU$^k$ networks on bounded Lipschitz domains with the same rate.","I.i.d. uniform collocation points produce a near-optimal high-probability residual estimate, up to a logarithmic factor and an arbitrarily small smoothness loss.","The elliptic norm equivalence transfers residual estimates directly to solution errors: $\\|u-u_{n,m}\\|_{H^{s+\\beta}}$ is comparable to the residual norm $\\|f-L_\\beta u_{n,m}\\|_{H^s}$."],"supporting_citations":[{"why":"supplies the comparison network theorem with the $O(h^{r-s})$ error bound and the exact low-degree coefficient identity used in the proof of Theorem 4.","marker":"[49]"},{"why":"provides the localized Jacobi kernel estimate that controls the off-diagonal decay in the Bernstein inequality's diagonal-dominance argument.","marker":"[73]"},{"why":"provides the polynomial approximation and spherical harmonic results used to compare continuous and empirical residual norms.","marker":"[19]"},{"why":"gives the spherical Marcinkiewicz-Zygmund inequality used to control empirical norms of low-degree polynomial components.","marker":"[59]"},{"why":"supplies the discrete least-squares comparison-and-stability framework that the deterministic proof adapts to network spaces.","marker":"[16]"},{"why":"underlies the homogeneous lifting construction used for the order-zero affine-network interpretation on bounded domains.","marker":"[4]"}],"fun_headline_variants":["Deterministic points match optimal rates in neural least squares","m ≈ n deterministic samples suffice for optimal elliptic approximation","No oversampling: deterministic collocation hits optimal rates","ReLU^k networks: discrete least squares achieves continuous-order accuracy","Optimal neural approximation from deterministic samples only"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire deterministic rate rests on a theorem from the authors' earlier work asserting that a target with the right parity can be approximated by a network whose low-frequency spherical coefficients match the target exactly; if that exact matching fails, the residual identity used in the proof breaks and the optimal-rate conclusion does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Deterministic points match optimal rates in neural least squares","m ≈ n deterministic samples suffice for optimal elliptic approximation","No oversampling: deterministic collocation hits optimal rates","ReLU^k networks: discrete least squares achieves continuous-order accuracy","Optimal neural approximation from deterministic samples only"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001124,"raw_usage":{"total_tokens":4828,"prompt_tokens":1249,"completion_tokens":3579,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":865,"completion_tokens_details":{"reasoning_tokens":3502}},"tokens_in":865,"tokens_out":3579,"duration_ms":23964,"temperature":1.0,"reasoning_tokens":3502,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:26:17.916429+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the construction behind the comparison-network theorem: compute the low-degree spherical harmonic coefficients of the comparison network and compare them with the target's coefficients for all degrees up to $C_1 h^{-1}$. If the identity $\\widehat{u_n}(\\nu,\\ell)=\\widehat{u}(\\nu,\\ell)$ fails at any degree in that range, then the proof's step $P_q L_\\beta u_n = P_q f$ no longer holds and the claimed rate is unsupported. An independent check would run the discrete least-squares problem with $m$ proportional to $n$ on a smooth parity-compatible right-hand side and test whether the empirical error decays at the predicted $n^{-(r-s)/d}$ rate over a wide range of $n$.","supporting_citations":[{"cited_title":"Petrushev and Y","cited_arxiv_id":null,"evidence_quote":"provides the localized Jacobi kernel estimate that controls the off-diagonal decay in the Bernstein inequality's diagonal-dominance argument."},{"cited_title":"Mhaskar, F","cited_arxiv_id":null,"evidence_quote":"gives the spherical Marcinkiewicz-Zygmund inequality used to control empirical norms of low-degree polynomial components."}],"review_version":1}