REVIEW 2 major objections 8 minor 24 references
Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method
T0 review · 2 major / 8 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read Harmonic functions with Barron boundary data are not Barron, yet can be approximated by low-norm Barron functions with only logarithmic norm growth.
desk verdict Core regularity result is clean and new; Deep Ritz rates have a fixable scaling inconsistency in the printed theorem. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
An explicit half-space harmonic corrector built from arctan and log, extended to rectangles by mirror charges plus a smooth cuboid eigenfunction corrector; the same family, regularized by a small vertical shift or by a smoothed logarithm, supplies the Barron approximants of logarithmic norm.
What would settle it
Compute the harmonic function on the unit square with boundary data equal to a single ReLU ridge function that crosses one side; check whether its gradient remains bounded near the kink or whether its second derivatives lie in L²—if either holds, the negative regularity claim is false.
Extended reading notes
Core claim
If the Dirichlet data g and a particular right-hand side generator U both lie in Barron space on a two-dimensional rectangle, the unique weak solution u* of the Poisson problem is generally neither Lipschitz nor in H² (hence not Barron), yet for every ε>0 there exist Barron approximants of norm O(|log ε|) that match either the boundary condition or the PDE exactly and converge to u* at rate essentially O(ε) in L∞, W^{1,q} and W^{2,p} for p<2.
Load-bearing premise
The constructive approximation rates and the Deep Ritz error bounds are proved only for half-spaces and for rectangles whose sides align with the coordinate axes; the underlying eigenfunction regularity fails as soon as the domain is a non-rectangular polygon or the operator is rotated.
Editorial extensions
If this is right
- Deep Ritz solvers with weight-decay regularization on rectangular domains achieve an H¹ error that decays like (log m)/√m when the number of neurons m tends to infinity.
- Exact representation of harmonic functions by bounded-weight shallow or deep ReLU networks is impossible once corners or kinks are present; only approximation with slowly growing norms is possible.
- The same logarithmic-norm approximants supply quantitative rates for physics-informed networks that use smoother activations whose second derivatives remain measures.
- Barron boundary data on a rectangle are completely characterized by the four one-dimensional Barron traces meeting continuously at the corners.
Reading between the lines
- The logarithmic barrier suggests that any neural architecture whose norm controls the Lipschitz constant will face the same mild obstruction on domains with corners.
- Extending the mirror-charge construction to polyhedral domains in higher dimensions would immediately give analogous rates for a much larger class of engineering geometries.
- The fact that the energy landscape has exponentially decaying tails but no minimizer may explain why gradient methods still succeed on these problems despite the absence of an exact Barron solution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies the Poisson/Dirichlet problem when the data live in the (representational) Barron space for shallow ReLU networks. Theorem 1 shows that on a 2D rectangle the weak solution with Barron boundary data is generally not Lipschitz, not in H², and not Barron — even for harmonic solutions with single-neuron boundary data σ(w·x+b) — yet it can be approximated to accuracy ε in L∞, W^{1,q} (q<∞) and W^{2,p} (p<2) by Barron functions of norm ≲ |log ε|, either matching the boundary data exactly (part 3) or the PDE exactly (part 4). Theorem 3 characterizes Barron boundary data on rectangles via one-dimensional Barron conditions on the sides. Theorem 2 applies the approximation theory to derive an a priori H¹ error estimate of order √(log m/√m) for a discretized, weight-decay-regularized Deep Ritz method. The proofs rest on an explicit closed-form harmonic extension on the half-space (Lemma 4), quantitative Barron-norm control of a logarithmic corrector via H²–H³ interpolation (Lemma 6), a mirror-charge reduction on quadrants (§3.2), and a smooth corrector on rectangles using cuboid eigenfunction regularity (Lemmas A.5–A.6), followed by weak-convergence passage to general Barron data.
Significance. If correct, this is a noteworthy contribution: it identifies a sharp and somewhat surprising failure of Barron-space regularity for the simplest elliptic boundary value problem — boundary effects destroy Lipschitz/H²/Barron regularity even for single-neuron data — while simultaneously showing the obstruction is quantitatively mild (|log ε| norm growth for ε-accuracy, i.e. exponentially decaying energy tails). The strengths are concrete: the constructions are explicit and closed-form (Eq. (3.1), (3.3), (3.8)), the norm bounds are derived by direct integration rather than fitting, the negative results (non-Lipschitz via Schwarz reflection; non-H² via the 1_{(0,∞)} ∉ H^{1/2} trace argument) are clean and rigorous, and the geometric scope limitation (half-spaces in any d, rectangles in d=2; failure of Lemma A.5 for non-rectangular polygons and misaligned operators) is openly flagged with a correct counterexample discussion after Lemma A.5. The Deep Ritz application is a natural and useful consequence. The results are checkable line-by-line and the paper is honestly written about what it does not prove (no lower bounds, aspect-ratio dependence deferred).
major comments (2)
- [§4, Theorem 2 statement and Steps 3–5a, Eq. (4.1)] As printed, the regularization scaling in Theorem 2 is inconsistent with the claimed rate. The theorem sets µ_m = γ√m with γ ≥ (∥g∥_B+1)/√m, i.e. µ_m ≥ ∥g∥_B+1 = O(1). But the competitor bound (4.1) contains the weight-decay contribution +Cµ∥g∥_B log m, which under this scaling is ≥ C(∥g∥_B+1)∥g∥_B log m — larger than the entire asserted bound C√(log 1/δ)(∥g∥²_B+1) log m/√m by a factor √m. Step 5a then asserts E_λ(u(a,W,b)) ≤ E_Dir(u*) + C√(log 1/δ)(∥g∥²_B+1) log m/√m, silently dropping this term. Hence no admissible γ yields the claimed H¹ rate as printed. Every other step of the proof is instead consistent with the alternative scaling µ_m = γ/√m, γ ≥ ∥g∥_B+1: (i) the Step 4 absorption condition ∥g∥_B log m/γ ≤ log m then requires exactly γ ≥ ∥g∥_B; (ii) the a priori bound ∥a∥²+∥W∥²+∥b∥² ≤ Ê/µ ≤ C√m matches the text (with the printed scaling it would be C/√m); (iii) the upper-branch exc
- [Appendix A, Lemma A.5 and its use in Lemma 10] The proof of Lemma A.5 asserts that Σ_n λ_n^k α_n(v)² defines an equivalent norm on H^k(Ω) for arbitrary v ∈ H^k(Ω). As stated this needs qualification: the sine basis diagonalizes the Dirichlet Laplacian, and the spectral characterization of H^k norms is standard only on the appropriate form domains (for k ≥ 1 this encodes boundary compatibility conditions; a general H¹ function on the cuboid does not have a sine series converging in H¹). The application in Lemma 10 — to w = u♯ − bû − V ∈ H¹_0(Ω) with ∆w = −∆V ∈ H^{k} — is the setting where the spectral argument is legitimate, so the result as used appears sound, but the statement and proof of Lemma A.5 should be made precise (e.g. by phrasing the equivalence for the Dirichlet realization of −∆ and its powers, or citing a source for the cuboid case). Since Lemma A.5 carries the higher-regularity step of the rectangle corrector and the a
minor comments (8)
- [§4, Step 2] The citations [LSSS14, Theorem 26.12], [LSSS14, Theorem 10.3] and [LSSS14, Theorem 6.8] appear to point to the wrong reference: LSSS14 is 'On the computational efficiency of training neural networks' (Livni–Shalev-Shwartz–Shamir), which contains no such theorems. The Rademacher generalization bound, the boosting bound for intersections of classes, and the fundamental theorem of PAC learning with these theorem numbers are in Shalev-Shwartz & Ben-David [SSBD14] (which the manuscript itself cites correctly elsewhere, e.g. Lemma 26.2). Please correct.
- [§1, Theorem 1(4)] In the statement of part (4), '∆u_ε ≡ ∆U' should read ∆ũ_ε ≡ ∆U. Also the notation ˜u_ε ∈ B ∩ (U + H²(Ω)) is slightly confusing since U is merely Barron; consider spelling out that ˜u_ε − U ∈ H²(Ω).
- [§4, Step 3] The bound ∥∇(u*−u_ε)∥²_{L²(Ω)} ≤ C∥g∥_B ε has an inconsistent (linear) dependence on ∥g∥_B; from Theorem 1(3) with q=2 one gets quadratic dependence, ∥g∥²_B ε² (which is stronger for small ε). The displayed claim is only used as an upper bound so nothing breaks, but it should be corrected for consistency.
- [§2.2, compact embedding bullet] The assertion ∇u ∈ BV(Ω) for Barron u is cited to [R W26, Lemma 1], a manuscript listed as 'in preparation, 2026'. Since this is used (if only for context), please either include a short proof (as was done for compact embedding in Lemma A.1) or provide a published reference.
- [§2.1] The review of polynomial approximation characterizations of Hölder/Sobolev spaces (including the one-dimensional induction) does not appear to be used later; the analogy to Barron spaces is only heuristic. Consider shortening this subsection to the Bramble–Hilbert statement actually needed.
- [Figures 1–3] The author notes the color scales differ across Figures 1–3; a common scale (or explicit colorbars) would make the mirror-charge construction easier to compare visually. In Figure 4 the finite-difference solution is a nice sanity check — a brief note on the discretization used would help.
- [§3.1, Lemma 6] The identification of u_ε with a Barron function u'_ε on Ω (via cut-off of the superlinearly growing v_ε) is correct but easy to miss; the sentence 'we will not distinguish between u_ε and u'_ε' would benefit from a forward pointer to Lemma A.2/A.3 where the quantitative norm bound for the cut-off is proved.
- [Throughout] Minor typos/typesetting: the plural 'perceptra' is nonstandard (it also appears in the [PPW23] title, where it may be intentional); in the Abstract 'Lebesgue and Sobolev norms (with at most two derivatives)' could be more precise; Eq. (3.2) the limit evaluation x1·sign(x1) is written in a slightly ambiguous inline form.
Circularity Check
No circularity: constructive potential-theory proofs with independent Barron-norm estimates; self-citations are background lemmas only.
full rationale
Theorem 1 is derived from an explicit half-space harmonic function (3.1), direct differentiation and integral estimates for W^{1,q}/W^{2,p} norms, a logarithmic corrector whose Barron norm is bounded by H^2–H^3 interpolation (s_ε=1/|log ε|), mirror-charge extension to quadrants, and a classical cuboid eigenfunction regularity lemma (A.5) plus boundary extension (A.6). None of these steps define the output in terms of itself, fit a free parameter to the target quantity, or rest on an unverified self-citation uniqueness theorem that forces the claim. Self-citations (EW20b, EW20c, CPV20, VW24) supply standard Barron embeddings, structure theorems, and a Liouville-type growth uniqueness used as ordinary lemmas; the target non-membership and approximation rates are proved by direct construction and contradiction (Schwarz reflection + 1_{(0,∞)}∉H^{1/2}), not imported. Theorem 2’s a priori rates build on Theorem 1 plus Rademacher/VC generalization; any scaling inconsistency in µ_m is a correctness gap, not a circular reduction. Theorem 3 is a direct one-dimensional characterization. The paper is self-contained against external benchmarks.
Assumptions & free parameters
assumptions (5)
- domain assumption Barron unit ball has finite Rademacher complexity and admits Maurey–Barron–Jones approximation rates in L²/H¹/L^∞ (EMW19a, EMWW20, Bac17).
- standard math Trace and extension theorems for Lipschitz domains; H^{1/2} characterization of boundary traces (Leo17, Dob10).
- standard math Dirichlet Laplacian on a cuboid admits a separable sine eigenbasis giving H^{k+2} regularity for H^k data (Lemma A.5).
- domain assumption H^{d/2+1+s} embeds (locally) into Barron space with explicit constant (CPV20 / Lemma A.2).
- standard math Schwarz reflection and Liouville-type uniqueness for subquadratically growing harmonics on half-spaces/quadrants.
Cite this review
Pith. "Pith review of Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method." pith.science (2026). https://pith.science/paper/VZXOLTMW
@misc{pith2026260725100,
author = {Pith},
title = {Pith review of: Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method},
year = {2026},
howpublished = {\url{https://pith.science/paper/VZXOLTMW}},
note = {Machine review of arXiv:2607.25100}
}
abstract
We prove that harmonic functions with Dirichlet boundary data in Barron space, a function class tailored to wide ReLU networks with a single hidden layer and suitably bounded weights, are generally neither Lipschitz continuous nor in the Sobolev class $H^2$. A fortiori, they are not in any function class in which the norm controls the Lipschitz constant, which rules out not only Barron space regularity, but also regularity in function classes for deeper ReLU networks with bounded coefficients. They can, however, be approximated to accuracy $\sim \varepsilon$ by Barron functions of low norm $\sim |\log\varepsilon|$ in various Lebesgue and Sobolev norms (with at most two derivatives). The positive result holds on very simple domains: Half-spaces in arbitrary dimension and rectangular domains in two dimensions. As an application of this regularity theory, we obtain a priori error estimates for Deep Ritz neural PDE solvers.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[5]
[DNPS23] Ronald DeVore, Robert D Nowak, Rahul Parhi, and Jonathan W Siegel. Weighted variation spaces and approximation by shallow relu networks.arXiv preprint arXiv:2307.15772,
-
[7]
Some observations on partial differential equations with Barron data
[EW20c] Weinan E and Stephan Wojtowytsch. Some observations on partial differential equations with Barron data. arXiv preprint arXiv:2012.01484,
arXiv 2012
-
[9]
Neural operator: Learning maps between function spaces.arXiv preprint arXiv:2108.08481,
36 STEPHAN WOJTOWYTSCH [KLL+21] Nikola Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew Stu- art, and Anima Anandkumar. Neural operator: Learning maps between function spaces.arXiv preprint arXiv:2108.08481,
-
[11]
[LJK19] Lu Lu, Pengzhan Jin, and George Em Karniadakis. Deeponet: Learning nonlinear operators for identi- fying differential equations based on the universal approximation theorem of operators.arXiv preprint arXiv:1910.03193,
arXiv 1910
-
[13]
[LMW20] Zhong Li, Chao Ma, and Lei Wu. Complexity measures for neural networks with general activation functions using path-based norms.arXiv preprint arXiv:2009.06132,
arXiv 2009
-
[17]
[SX21a] Jonathan W Siegel and Jinchao Xu. Characterization of the variation spaces corresponding to shallow neural networks.arXiv preprint arXiv:2106.15002,
-
[19]
[SX21c] Jonathan W Siegel and Jinchao Xu. Sharp bounds on the approximation rates, metric entropy, andn-widths of shallow neural networks.arXiv preprint arXiv:2101.12365,
-
[21]
[VW24] Malhar Vaishampayan and Stephan Wojtowytsch. Solving the poisson equation with dirichlet data by shallow relu α-networks: A regularity and approximation perspective.arXiv preprint arXiv:2412.07728,
Show all 24 references
-
[22]
On the global convergence of gradient descent training for two-layer Relu networks in the mean field regime.arXiv:2005.13530 [math.AP],
[Woj20] Stephan Wojtowytsch. On the global convergence of gradient descent training for two-layer Relu networks in the mean field regime.arXiv:2005.13530 [math.AP],
2005 arXiv
-
[23]
Optimal bump functions for shallow ReLU networks: Weight decay, depth separation and the curse of dimensionality.arXiv:2209.01173 [stat.ML],
[Woj22] Stephan Wojtowytsch. Optimal bump functions for shallow ReLU networks: Weight decay, depth separation and the curse of dimensionality.arXiv:2209.01173 [stat.ML],
-
[24]
A deep learning framework for multi- operator learning: Architectures and approximation theory.arXiv preprint arXiv:2510.25379,
[WSZS25] Adrien Weihs, Jingmin Sun, Zecheng Zhang, and Hayden Schaeffer. A deep learning framework for multi- operator learning: Architectures and approximation theory.arXiv preprint arXiv:2510.25379,
-
[1967]
On the approximation properties of neural networks.arXiv preprint arXiv:1904.02311,
[SX19] Jonathan W Siegel and Jinchao Xu. On the approximation properties of neural networks.arXiv preprint arXiv:1904.02311,
1904 arXiv
-
[1993]
Optimal learning.arXiv preprint arXiv:2203.15994,
[BBDP22] Peter Binev, Andrea Bonito, Ronald DeVore, and Guergana Petrova. Optimal learning.arXiv preprint arXiv:2203.15994,
-
[1996]
A function space view of bounded norm infinite width ReLU nets: The multivariate case.arXiv preprint arXiv:1910.01635,
[OWSS19] Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro. A function space view of bounded norm infinite width ReLU nets: The multivariate case.arXiv preprint arXiv:1910.01635,
1910 arXiv
-
[1999]
Group equi- variant fourier neural operators for partial differential equations.arXiv preprint arXiv:2306.05697,
[HZF+23] Jacob Helwig, Xuan Zhang, Cong Fu, Jerry Kurtin, Stephan Wojtowytsch, and Shuiwang Ji. Group equi- variant fourier neural operators for partial differential equations.arXiv preprint arXiv:2306.05697,
-
[2014]
Neural scaling laws of deep relu and deep operator network: A theoretical study.arXiv preprint arXiv:2410.00357,
[LZLS24] Hao Liu, Zecheng Zhang, Wenjing Liao, and Hayden Schaeffer. Neural scaling laws of deep relu and deep operator network: A theoretical study.arXiv preprint arXiv:2410.00357,
-
[2015]
A priori estimates of the population risk for two-layer neural networks
[EMW18] Weinan E, Chao Ma, and Lei Wu. A priori estimates of the population risk for two-layer neural networks. Comm. Math. Sci., 17(5):1407 – 1425 (2019), arxiv:1810.06397 [cs.LG] (2018). [EMW19a] Weinan E, Chao Ma, and Lei Wu. The Barron space and the flow-induced function s...
2019 arXiv
-
[2019]
Fourier neural operator for parametric partial differential equations.arXiv preprint arXiv:2010.08895,
[LKA+20] Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stu- art, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations.arXiv preprint arXiv:2010.08895,
2010 arXiv
-
[2020]
Uniform convergence guarantees for the deep ritz method for nonlinear problems.Advances in Continuous and Discrete Models, 2022(1):49,
[DMZ22] Patrick Dondl, Johannes M¨ uller, and Marius Zeinhofer. Uniform convergence guarantees for the deep ritz method for nonlinear problems.Advances in Continuous and Discrete Models, 2022(1):49,
2022
-
[2021]
Machine learning for ellip- tic PDEs: Fast rate generalization bound, neural scaling law and minimax optimality.arXiv preprint arXiv:2110.06897,
[LCL+21] Yiping Lu, Haoxuan Chen, Jianfeng Lu, Lexing Ying, and Jose Blanchet. Machine learning for ellip- tic PDEs: Fast rate generalization bound, neural scaling law and minimax optimality.arXiv preprint arXiv:2110.06897,
-
[2022]
Penalising the biases in norm regularisation enforces sparsity
[BF23] Etienne Boursier and Nicolas Flammarion. Penalising the biases in norm regularisation enforces sparsity. arXiv preprint arXiv:2303.01353,
-
[2023]
Neural network approximation and estimation of classifiers with classification boundary in a Barron class.arXiv:2011.09363 [math.F A],
[CPV20] Andrei Caragea, Philipp Petersen, and Felix Voigtlaender. Neural network approximation and estimation of classifiers with classification boundary in a Barron class.arXiv:2011.09363 [math.F A],
2011 arXiv
-
[2024]
[Voi22] Felix Voigtlaender.l p-sampling numbers for the Fourier-analytic Barron space.arXiv preprint arXiv:2208.07605,
-
[2025]
Embedding inequalities for barron-type spaces.arXiv preprint arXiv:2305.19082,
[Wu23] Lei Wu. Embedding inequalities for barron-type spaces.arXiv preprint arXiv:2305.19082,
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.