Pith. sign in

REVIEW 3 major objections 4 minor 48 references

The Barron-Lipschitz Energy Gap and Depth Separation Phenomena in Scientific Machine Learning

T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Infinite-width shallow neural networks provably miss the optimal thin-sheet fold because their two-dimensional structure allows only straight folds; a two-layer composition can fold along a circle.

desk verdict The Barron-Lipschitz gap is a real idea and Theorem 3 is solid, but Theorem 7 as written has an arithmetic slip and the structure theorem underpinning it is only sketched, so treat the shell example as conditional. read the letter →

arxiv 2607.25905 v1 pith:OCNWHQG5 submitted 2026-07-28 math.AP stat.ML

classification math.APstat.ML MSC 46E3549Q2068T0774K20
keywords BarronspaceenergygapdepthseparationthinelasticshellcalculusofvariationsLipschitzfunctionsstructuretheoremscientificmachinelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Shallow neural networks of infinite width—the class called Barron functions—are shown to have an energy gap in a model of thin elastic sheets: a circular fold lowers the energy, but Barron functions can only create folds along straight lines, so they miss the minimizer. A composition of two Barron functions produces the circular fold, making the gap a depth-separation phenomenon rather than a smoothness artifact. The key step is a two-dimensional structure theorem: a Barron function whose Hessian is square-integrable outside a singular set has that singular set concentrated on a countable union of straight lines. The paper also proves there is no Barron-Lipschitz gap for a large class of integral first-order functionals, and exhibits one-dimensional functionals with singular integrands or L∞ constraints where the gap is positive.

What carries the argument

The load-bearing machinery is the two-dimensional structure theorem for Barron functions (Theorem 2): if f is Barron and its distributional Hessian is square-integrable outside a relatively closed set K of locally finite H¹ measure, then the singular part of the Hessian is supported on a countable, locally finite union of straight lines contained in K. This forces Barron folds to run along entire straight lines, which is geometrically incompatible with folding along the mid-circle of an annulus. The matching upper bound is built by composing two Barron functions: one computes a radial coordinate whose non-smooth level set is a circle, and the other applies a V-shaped one-dimensional profile,

What would settle it

Exhibit a Barron function f on R² and a relatively closed set K of finite H¹ measure such that D²f∈L²(Ω\K) but the singular support of the Hessian contains a smooth curve that is not a straight line, e.g., a circular arc. Such a counterexample would falsify Theorem 2 and remove the lower bound on Barron energies in Theorem 7.

Watch

Extended reading notes

Core claim

The central discovery is a quantitative energy gap between Barron functions and more flexible function classes for a linearized model of bending, stretching, and folding of a thin sheet. On an annulus with anchored boundary and a target metric favoring radial stretching, the infimum of the energy over general admissible pairs is strictly smaller than the infimum over Barron functions; the theorem states a gap of at least 2π/49 for parameter ranges it makes explicit. The better competitor is a composition of two Barron functions whose non-smooth set is a circle. The reason Barron functions cannot match it is the structure theorem: their singular Hessian set must be a locally finite union of s

Load-bearing premise

The thin-sheet gap rests on the two-dimensional structure theorem that a Barron function's singular Hessian set is a countable union of straight lines; if that theorem fails (and its appendix proof is compressed), the circular-fold gap has no basis.

Editorial extensions

If this is right

  • For continuous integral first-order functionals on bounded finite-perimeter domains, the infimum over Barron functions equals the infimum over Lipschitz functions, so shallow networks introduce no Lavrentiev-type gap in those settings.
  • In one dimension, functionals with an L∞ constraint or an integrand singular near the boundary can have a positive Barron-Lipschitz gap; for some of these the near-minimizers themselves differ in shape between the two classes.
  • For the thin-sheet model with anchored boundary on an annulus, the energy gap is at least 2π/49 and is realized by a two-layer composition, establishing a depth-separation phenomenon in scientific machine learning.
  • Numerical solvers built on shallow ReLU networks inherit the straight-fold restriction and may systematically overestimate the minimal energy in geometric variational problems with curved crease patterns.
  • The no-gap result extends to vector-valued functions, boundary terms, and more general integrands under the technical conditions given in the paper, so the positive results cover a broad family of variational settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct numerical test could train a shallow ReLU network and a two-layer network on the annular thin-sheet energy; if the predicted gap holds, the shallow network should consistently settle at a strictly higher energy, near the theorem's gap constant.
  • The straight-line structure theorem suggests a design principle: any variational problem whose low-energy competitors require non-flat singular sets will be structurally constrained for shallow ReLU ansatzes; architectures that compute radial or higher-order first-layer features may bypass the gap.
  • The one-dimensional examples indicate that the practical impact of a Barron-Lipschitz gap depends on whether the shape of near-minimizers changes, not just the energy value; some gaps are only in the energy, while others change the minimizing morphology.
  • The circular-fold gap may extend to higher dimensional hyper-surfaces or non-radial domains with multiple folds, but the paper proves it only in the annular case; a similar analysis for perturbed or non-radial domains would be a natural next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies variational energy gaps between Barron functions (infinite-width ReLU networks) and broader classes such as Lipschitz functions or H^2 functions outside a singular set. Theorem 3 shows that for a large class of integral first-order functionals, the infimum over Barron functions equals that over Lipschitz functions. Theorems 5 and 6 construct one-dimensional first-order functionals exhibiting a Barron-Lipschitz gap, the second with a sketch. The main result is Theorem 7: for a model thin-shell energy on an annulus, a two-layer composition of Barron functions achieves lower energy than any Barron function, with an explicit quantitative gap of 2π/49. This depth-separation result rests on the new Theorem 2, which asserts that in two dimensions the non-L^2 part of the Hessian of a Barron function is supported on a countable, locally finite union of lines.

Significance. If the main results hold, the paper makes a substantive contribution to the interface of Barron space theory and the calculus of variations. The positive result Theorem 3 is clean and useful: it shows that for many standard first-order integral functionals there is no additional Lavrentiev-type gap caused by the Barron constraint. Theorem 7 is conceptually striking: it gives a quantitative variational setting in which shallow networks are provably inferior to deeper compositions, with the energy difference a fixed constant independent of width. The structure theorem Theorem 2, if fully established, would be an important regularity result for two-dimensional Barron functions. However, the proof of Theorem 2 is only sketched, and the proof of Theorem 7 contains a concrete algebraic error in the central inequality. The overall claims are plausible and likely repairable, but the manuscript in its current form does not yet provide a complete proof of its headline result.

major comments (3)
  1. [§5.4, Step 2 of proof of Theorem 7] The displayed chain of inequalities contains an algebraic error. From the preceding line one obtains E ≥ H¹(B)(inf F + 1) + H¹(G)(inf F − 1/7) = 2π inf F + H¹(B) − H¹(G)/7. The manuscript writes this as 2π inf F + (H¹(B) − H¹(G))/7, which is not an equality unless H¹(B)=0. With the correct expression, 2π inf F + (8/7)H¹(B) − 2π/7, the bound H¹(B) ≥ 2π/7 from Lemma 10 gives exactly the claimed gap +2π/49. Thus the conclusion is recoverable, but the proof as written does not establish the stated inequality; a corrected derivation must replace the erroneous display.
  2. [Appendix A, proof of Theorem 2] The proof is a compressed sketch of a load-bearing structural result. The decomposition of the Barron spectral measure into atoms, great-circle parts, and a remainder is asserted rather than proved; the classification of blow-up limits as piecewise linear is not fully justified; and the claim that a non-piecewise-linear positively one-homogeneous limit would force D²f ∉ L²(Ω\K) is only supported by a scaling computation that does not by itself distinguish piecewise-linear from non-piecewise-linear behavior. Since Theorem 7's lower bound for Barron functions depends directly on this theorem, a complete proof with all technical steps spelled out is necessary before the main claim can be considered established.
  3. [§5.4, Step 2, replacement of K by K'] The sentence "As this only decreases the energy" is not immediate: replacing K by a subset K' enlarges the domain of integration in the bending term. The correct justification is that K and K' differ by a set of zero Lebesgue measure (an H¹-finite set in R²), so the Lebesgue integral over Ω\K equals that over Ω\K', while the folding penalty decreases. The manuscript should state this explicitly; as written the monotonicity claim is misleading and could be read in the opposite direction.
minor comments (4)
  1. [Theorem 7 statement] The last condition in the maximum defining μ appears garbled: it reads like "emin − 1/(2r²(...)²)", which cannot be what is meant. It should presumably be the fraction (log(R/r)+λ(R+r)/2)/(2r²(1/7−log(R/r)−2λ(1−r/R)(2r+R)/9)²). Please fix the typesetting.
  2. [§2, fact (3) vs. proof of Theorem 3] The introduction cites [EW20b, Theorem 3.2] for the fact that every smooth function coincides with a Barron function on a bounded set, while the proof of Theorem 3 cites [EW20b, Theorem 3.1]. Please make the references consistent.
  3. [Proof of Lemma 10] The argument that any point of S¹ lies in at most two intervals after the pruning step is stated tersely and relies on a lifting of intervals from S¹ to R. This is plausible but should be written out, especially because the 'unique point' property is essential for the factor 2 in the estimate.
  4. [Theorem 5, Step 2] The deduction that lim_{t→0+} u'(t)=0 forces the L∞ term to be at least 1 uses the fact that the essential supremum is at least the limit superior of pointwise values of the precise representative. This is standard but should be stated explicitly for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the thin-shell gap uses an explicitly constructed competitor and an independent (if compressed) structure theorem; the flagged defect is an arithmetic gap in Theorem 7, not circularity.

full rationale

Walking the paper's derivation chain, no claim reduces to its own inputs. Theorem 3 proves density of Barron functions in the Lipschitz class by mollifying a Lipschitz function and invoking [EW20b] only for the standard fact that smooth compactly supported functions are Barron; the cited theorem's assumptions do not include the no-gap conclusion, so this is independent support rather than a fitted input called a prediction. The one-dimensional gaps (Theorems 5 and 6) are established directly from the one-dimensional characterization u′∈BV for Barron functions and by explicit sawtooth Lipschitz competitors; no parameter is fitted to the claimed gap. For Theorem 7, the energy competitor (u1∘u2, S) is constructed explicitly with energy bounded by 2π(log(R/r)+λ(R+r)/2), and the lower bound for Barron functions is argued through the Appendix structure theorem (Theorem 2) on the singular support of the Hessian. That structure theorem is independent of the elastic energy and is proved, though in compressed form, from the Barron spectral representation rather than imported as the desired energy-gap conclusion. Self-citations such as [EW20b], [Woj22], and [EW20a] supply background representation facts, not the theorem being proved, so they are not load-bearing in a circular way. The serious problem visible in the proof is arithmetic rather than circularity: with H1(B)≥2π/7, the identity B+G=2π gives (H1(B)−H1(G))/7 = (2H1(B)−2π)/7 ≥ −10π/49, not the claimed +2π/49; obtaining the displayed gap would require H1(B)≥8π/7. That is a correctness risk, not a reduction of the conclusion to its inputs.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

All non-standard prerequisites are explicitly cited or proved. Theorem 7's parameters (r,R,λ,μ) are part of the statement's sufficient conditions, not fitted constants; the example (r,R,λ,μ)=(14,15,0.1,40) merely illustrates the inequalities.

assumptions (5)
  • domain assumption Barron space structural and approximation properties (representation formula, Lipschitz embedding, finite-network approximation) as established in EW20b.
    Used as foundation in Section 2 and throughout the proofs.
  • domain assumption Existence of a Barron function matching the boundary trace of w in Theorem 3(2).
    Explicit assumption in the theorem; without it the constrained infimum equality over B would not follow.
  • domain assumption The sliced/linearized elastic energy (5.2) is a valid toy model for thin-shell folding.
    Physical relevance of Theorem 7; the model is linearized and not the full 3D elasticity.
  • standard math Standard analytic tools: Rademacher's theorem, Kirszbraun extension, mollifier convergence, dominated convergence, direct method of the calculus of variations.
    Invoked in proofs of Theorems 3, 5, 6 and Lemma 8.
  • standard math The measure decomposition of the Barron spectral measure into atoms, great-circle-supported parts, and a remainder with no great-circle mass.
    Used in the proof of Theorem 2 (Appendix A).

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Barron-Lipschitz Energy Gap and Depth Separation Phenomena in Scientific Machine Learning." pith.science (2026). https://pith.science/paper/OCNWHQG5

@misc{pith2026260725905,
  author       = {Pith},
  title        = {Pith review of: The Barron-Lipschitz Energy Gap and Depth Separation Phenomena in Scientific Machine Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OCNWHQG5}},
  note         = {Machine review of arXiv:2607.25905}
}
read the original abstract

We illustrate in several examples that even neural networks of infinite width (specifically, Barron functions) may encounter substantial obstacles when used as a model class for problems in the calculus of variations. An instance of practical relevance concerns the bending, stretching and folding of a thin elastic shell with anchored or clamped boundary conditions where elastic energy could be reduced by folding along a circular line, but the neural networks can only describe straight folds along entire lines. Conversely, we show that there is no gap between the energy that Barron functions and Lipschitz functions can achieve for a large class of integral first-order functionals.

Figures

Figures reproduced from arXiv: 2607.25905 by the authors.

Figure 1
Figure 1. Top row: The sets A, B, C on the circle correspond to the angles on the unit circle for which the line intersects the inner, central and outer ring respectively. Depending on the distance of the line to the origin one (or multiple) of the sets may be empty, and they may have one or two connected components, see also [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. The set of ‘bad angles’ (red) and ‘good angles’ (green) for various line arrangements and different values of r and R. The ‘good’ set corresponds to the set B in [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. As the distance of the straight line from the origin varies, the sets A, B, C as in [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: As above, we observe that |B| = cos−1 x b  − cos−1 x a  |A| = cos−1  x R  − cos−1 x b  . Therefore we note that |B| |A| + |C| ≤ |B| |A| ≤ cos−1 ( x b ) − cos−1 ( x a ) cos−1( x R ) − cos−1( x b ) . Thus, we define h(x) := cos−1 ( x b ) − cos−1 ( x a ) cos−1( x …

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

48 extracted references · 8 linked inside Pith

  1. [1]

    Mathematical analysis, 1958

    Tom M Apostol and CM Ablow. Mathematical analysis, 1958

  2. [2]

    Functions of bounded variation and free discontinuity problems

    Luigi Ambrosio, Nicola Fusco, and Diego Pallara. Functions of bounded variation and free discontinuity problems . Oxford university press, 2000

  3. [3]

    Variational models for phase transitions, an approach via -convergence

    Giovanni Alberti. Variational models for phase transitions, an approach via -convergence. In Calculus of variations and partial differential equations: topics on geometrical evolution problems and degree theory , pages 95--114. Springer, 2000

  4. [4]

    Breaking the curse of dimensionality with convex neural networks

    Francis Bach. Breaking the curse of dimensionality with convex neural networks. Journal of Machine Learning Research , 18(19):1--53, 2017

  5. [5]

    Discontinuous equilibrium solutions and cavitation in nonlinear elasticity

    John Macleod Ball. Discontinuous equilibrium solutions and cavitation in nonlinear elasticity. Philosophical Transactions of the Royal Society of London. Series A, Mathematical and Physical Sciences , 306(1496):557--611, 1982

  6. [6]

    Onsager's conjecture for admissible weak solutions

    Tristan Buckmaster, Camillo De Lellis, L \'a szl \'o Sz \'e kelyhidi Jr, and Vlad Vicol. Onsager's conjecture for admissible weak solutions. arXiv preprint arXiv:1701.08678 , 2017

  7. [7]

    New examples on L avrentiev gap using fractals

    Anna Kh Balci, Lars Diening, and Mikhail Surnachev. New examples on L avrentiev gap using fractals. Calculus of Variations and Partial Differential Equations , 59(5):180, 2020

  8. [8]

    Penalising the biases in norm regularisation enforces sparsity

    Etienne Boursier and Nicolas Flammarion. Penalising the biases in norm regularisation enforces sparsity. arXiv preprint arXiv:2303.01353 , 2023

Show all 48 references
  1. [9]

    A global method for relaxation in w^ 1, p and in sbv^p

    Guy Bouchitt \'e , Irene Fonseca, Giovanni Leoni, and Lu \' sa Mascarenhas. A global method for relaxation in w^ 1, p and in sbv^p . Archive for rational mechanics and analysis , 165(3):187--242, 2002

  2. [10]

    Finite element methods for the stretching and bending of thin structures with folding

    Andrea Bonito, Diane Guignard, and Angelique Morvant. Finite element methods for the stretching and bending of thin structures with folding. arXiv preprint arXiv:2311.04810 , 2023

  3. [11]

    h-principle and rigidity for c^ 1, -isometric embeddings

    Sergio Conti, Camillo De Lellis, and L \'a szl \'o Sz \'e kelyhidi Jr. h-principle and rigidity for c^ 1, -isometric embeddings. In Nonlinear Partial Differential Equations: The Abel Symposium 2010 , pages 83--116. Springer, 2010

  4. [12]

    Deformation concentration for martensitic microstructures in the limit of low volume fraction

    Sergio Conti, Johannes Diermeier, and Barbara Zwicknagl. Deformation concentration for martensitic microstructures in the limit of low volume fraction. Calculus of Variations and Partial Differential Equations , 56(1):16, 2017

  5. [13]

    Onsager's conjecture on the energy conservation for solutions of E uler's equation

    Peter Constantin, Weinan E, and Edriss S Titi. Onsager's conjecture on the energy conservation for solutions of E uler's equation. Commun.Math. Phys. , 1994

  6. [14]

    On turbulence and geometry: from N ash to O nsager

    Camillo De Lellis and L \'a szl \'o Sz \'e kelyhidi Jr. On turbulence and geometry: from N ash to O nsager. arXiv preprint arXiv:1901.02318 , 1, 2019

  7. [15]

    Angewandte Funktionalanalysis: Funktionalanalysis, Sobolev-R \"a ume und elliptische Differentialgleichungen

    Manfred Dobrowolski. Angewandte Funktionalanalysis: Funktionalanalysis, Sobolev-R \"a ume und elliptische Differentialgleichungen . Springer-Verlag, 2010

  8. [16]

    Measure theory and fine properties of functions

    Lawrence C Evans and Ronald F Gariepy. Measure theory and fine properties of functions . Chapman and Hall/CRC, 2025

  9. [17]

    The B arron space and the flow-induced function spaces for neural network models

    Weinan E, Chao Ma, and Lei Wu. The B arron space and the flow-induced function spaces for neural network models. Constructive Approximation , 55(1):369--406, 2022

  10. [18]

    L’h \^o pital’s monotone rule, gromov’s theorem, and operations that preserve the monotonicity of quotients

    Ricardo Estrada and Miroslav Pavlovi \'c . L’h \^o pital’s monotone rule, gromov’s theorem, and operations that preserve the monotonicity of quotients. Publications de l'Institut Mathematique , 101(115):11--24, 2017

  11. [19]

    The power of depth for feedforward neural networks

    Ronen Eldan and Ohad Shamir. The power of depth for feedforward neural networks. In Conference on learning theory , pages 907--940. PMLR, 2016

  12. [20]

    Partial differential equations , volume 19

    Lawrence C Evans. Partial differential equations , volume 19. American mathematical society, 2022

  13. [21]

    On the banach spaces associated with multi-layer relu networks: Function representation, approximation theory and gradient descent dynamics

    Weinan E and Stephan Wojtowytsch. On the banach spaces associated with multi-layer relu networks: Function representation, approximation theory and gradient descent dynamics. arXiv preprint arXiv:2007.15623 , 2020

  14. [22]

    Representation formulas and pointwise properties for B arron functions

    Weinan E and Stephan Wojtowytsch. Representation formulas and pointwise properties for B arron functions. Calc. Var. Partial Differential Equations , 61(46), 2020

  15. [23]

    The L avrentiev gap phenomenon in nonlinear elasticity

    M Foss, WJ Hrusa, and VJ Mizel. The L avrentiev gap phenomenon in nonlinear elasticity. Archive for rational mechanics and analysis , 167:337--365, 2003

  16. [24]

    A hierarchy of plate models derived from nonlinear elasticity by gamma-convergence

    Gero Friesecke, Richard D James, and Stefan M \"u ller. A hierarchy of plate models derived from nonlinear elasticity by gamma-convergence. Archive for rational mechanics and analysis , 180(2):183--236, 2006

  17. [25]

    An anisotropic P oincar \'e inequality in gsbv^p and the limit of strongly anisotropic Mumford--Shah functionals

    Janusz Ginster and Peter Gladbach. An anisotropic P oincar \'e inequality in gsbv^p and the limit of strongly anisotropic Mumford--Shah functionals. ESAIM: Control, Optimisation and Calculus of Variations , 32:23, 2026

  18. [26]

    Minimal surfaces and functions of bounded variation

    Enrico Giusti. Minimal surfaces and functions of bounded variation . Birkh\"auser, 1984

  19. [27]

    Elliptic partial differential equations of second order , volume 224

    David Gilbarg and Neil S Trudinger. Elliptic partial differential equations of second order , volume 224. springer, 2015

  20. [28]

    Sur quelques problemes du calcul des variations

    Mikhail Lavrentieff. Sur quelques problemes du calcul des variations. Annali di Matematica Pura ed Applicata , 4(1):7--28, 1927

  21. [29]

    On the L avrentiev phenomenon

    Philip D Loewen. On the L avrentiev phenomenon. Canadian Mathematical Bulletin , 30(1):102--108, 1987

  22. [30]

    Strong approximation of special functions of bounded variation functions with prescribed jump direction

    Giuliano Lazzaroni, Piotr Wozniak, and Caterina Ida Zeppieri. Strong approximation of special functions of bounded variation functions with prescribed jump direction. Mathematische Nachrichten , 298(1):312--327, 2025

  23. [31]

    Sets of finite perimeter and geometric variational problems: an introduction to Geometric Measure Theory

    Francesco Maggi. Sets of finite perimeter and geometric variational problems: an introduction to Geometric Measure Theory . Cambridge University Press, 2012

  24. [32]

    B. Mani\`a. Sopra un esempio di L avrentieff. Boll. Un. Mat. Ital. B , 13, 1934

  25. [33]

    The L avrentiev gap phenomenon for harmonic maps into spheres holds on a dense set of zero degree boundary data

    Katarzyna Mazowiecka and Pawe Strzelecki. The L avrentiev gap phenomenon for harmonic maps into spheres holds on a dense set of zero degree boundary data. Advances in Calculus of Variations , 10(3):303--314, 2017

  26. [34]

    C ^1 -isometric imbeddings

    John Nash. C ^1 -isometric imbeddings. Annals of mathematics , 60(3):383--396, 1954

  27. [35]

    The imbedding problem for R iemannian manifolds

    John Nash. The imbedding problem for R iemannian manifolds. Annals of mathematics , 63(1):20--63, 1956

  28. [36]

    A function space view of bounded norm infinite width relu nets: The multivariate case

    Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro. A function space view of bounded norm infinite width relu nets: The multivariate case. arXiv preprint arXiv:1910.01635 , 2019

  29. [37]

    Banach space representer theorems for neural networks and ridge splines

    Rahul Parhi and Robert D Nowak. Banach space representer theorems for neural networks and ridge splines. Journal of Machine Learning Research , 22(43):1--40, 2021

  30. [38]

    Minimum norm interpolation by perceptra: Explicit regularization and implicit bias

    Jiyoung Park, Ian Pelakh, and Stephan Wojtowytsch. Minimum norm interpolation by perceptra: Explicit regularization and implicit bias. Advances in Neural Information Processing Systems , 36:8714--8756, 2023

  31. [39]

    Real and complex analysis

    Walter Rudin. Real and complex analysis . McGraw-Hill, Inc., 1987

  32. [40]

    Benefits of depth in neural networks

    Matus Telgarsky. Benefits of depth in neural networks. In Conference on learning theory , pages 1517--1539. PMLR, 2016

  33. [41]

    Ridges, neural networks, and the radon transform

    Michael Unser. Ridges, neural networks, and the radon transform. Journal of Machine Learning Research , 24(37):1--33, 2023

  34. [42]

    Depth separation beyond radial functions

    Luca Venturi, Samy Jelassi, Tristan Ozuch, and Joan Bruna. Depth separation beyond radial functions. Journal of machine learning research , 23(122):1--56, 2022

  35. [43]

    Solving the Poisson Equation with Dirichlet data by shallow ReLU ^ -networks: A regularity and approximation perspective

    Malhar Vaishampayan and Stephan Wojtowytsch. Solving the Poisson Equation with Dirichlet data by shallow ReLU ^ -networks: A regularity and approximation perspective. arXiv preprint arXiv:2412.07728 , 2024

  36. [44]

    Lipschitz algebras

    Nik Weaver. Lipschitz algebras . World Scientific, 2018

  37. [45]

    Optimal bump functions for shallow ReLU networks: Weight decay, depth separation and the curse of dimensionality

    Stephan Wojtowytsch. Optimal bump functions for shallow ReLU networks: Weight decay, depth separation and the curse of dimensionality. arXiv preprint arXiv:2209.01173 , 2022

  38. [46]

    A note on elliptic regularity theory in barron spaces

    Stephan Wojtowytsch. A note on elliptic regularity theory in barron spaces. in preparation , 2026

  39. [47]

    Averaging of functionals of the calculus of variations and elasticity theory

    Vasilii Vasil'evich Zhikov. Averaging of functionals of the calculus of variations and elasticity theory. Mathematics of the USSR-Izvestiya , 29(1):33, 1987

  40. [48]

    Lavrentiev phenomenon and homogenization for some variational problems

    VV Zhikov. Lavrentiev phenomenon and homogenization for some variational problems. Composite media and homogenization theory , pages 273--288, 1995

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.