REVIEW 3 major objections 6 minor 43 references
Approximation Theory and Applications of Randomized Neural Networks for Solving High-Dimensional PDEs
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that randomized neural networks achieve dimension-independent Sobolev approximation rates, making them a practical tool for high-dimensional PDEs.
desk verdict Serious attempt at dimension-independent H^1/H^2 approximation rates for tanh RaNNs, but a false identity in Eq. (3.12) breaks the central proof; the numerics are plausible and the paper deserves major revision, not a full reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the smooth Heaviside approximation H_ε(x) = ½(1+tanh(x/ε)), introduced in equation (3.2). The proof starts from a representation of the target as an integral of a shifted Heaviside function, replaces the sharp step by H_ε, and then identifies the resulting expectation over random weights, biases, and shifts with a tanh randomized neural network. The approximation error is decomposed into the smoothing error from H vs H_ε and the Monte Carlo sampling error of a finite-width network, and both are bounded explicitly in the $H^{1}$ and $H^{2}$ norms.
What would settle it
Evaluate the unproved equality in equation (3.12) numerically for a smooth target such as u(x)=$e^{{-x^2}}$: if the difference between the two integrals does not decay polynomially in the smoothing parameter ε, the error decomposition in the proof is incomplete and the claimed rates do not follow.
Extended reading notes
Core claim
The central discovery is that a randomized neural network of width N with tanh activations approximates Sobolev functions in the $H^{2}$ norm on a ball with an error bounded by C $M^{2}$ $d^{2}$ ||u||_{$H^{{4+s}}$} $N^{{-(γ/(γ+2))(s-d/2)/(2s+3)}}$ for any γ in (0,1/2), and a similar power law in the $H^{1}$ norm for functions in $H^{{3+s}}$. The exponents are independent of the dimension d whenever the regularity s exceeds d/2, and they approach $N^{{-1/10}}$ in $H^{2}$ and $N^{{-1/4}}$ in $H^{1}$ as the regularity grows. The proof represents the target function as a Fourier integral, rewrites it using a smooth tanh-based approximation to the Heaviside step function, and then shows that the resulting expectation over random weights, biases, and shifts is exactly a randomized neural network; the finite-width network is a Monte Carlo estimate of that expectation, and the error splits into a smoothing error and a sampling error.
Load-bearing premise
The proof takes as given that replacing the sharp step function by its smooth tanh approximation and shifting the integration limits leaves no unaccounted boundary error, and the size of that residual is never bounded.
Editorial extensions
If this is right
- If the rates hold, a physics-informed extreme learning machine that solves a linear PDE by a single least-squares solve with N randomized tanh features has a total error governed by the approximation error, so the cost to reach a target accuracy in dimension d scales polynomially rather than exponentially in d.
- The H^2 approximation bound directly controls the PDE residual for second-order operators, since ||L[u_θ]||_{L2} is bounded by a constant times ||u_θ - u||_{H^2}; a small Sobolev error therefore implies a small physics-informed loss.
- For the heat equation, Black–Scholes, and Heston models in dimension up to 100, the numerical experiments report relative L2 errors of 1–3% with runtimes under six minutes, consistent with the claim that the curse of dimensionality is alleviated in practice.
- The dimension-independent rates imply that increasing the width N improves accuracy without needing to add features per dimension; the hidden weights are drawn once and fixed, and only the output layer is fitted.
Reading between the lines
- The same representation-to-Monte-Carlo pipeline could yield explicit approximation bounds for other activation functions that are scaled antiderivatives of localized bumps, such as the sigmoid, as long as the smoothing error can be controlled in Sobolev norms.
- The paper's rates come with constants that grow like d^2 and factors that depend on the Fourier integrability of the target; for functions whose Fourier transforms concentrate away from zero, these constants may grow exponentially in d, so the practical range of dimensions for a given accuracy may be smaller than the asymptotic result suggests.
- Because the H^2 norm controls first and second derivatives, the same dimension-independent rates should apply to financial Greeks (derivatives of option prices) obtained from the randomized-network solution of Black–Scholes or Heston models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies randomized neural networks (RaNNs) in the physics-informed extreme learning machine (PIELM) framework for high-dimensional PDEs. The theoretical part claims that tanh-activated RaNNs can approximate functions of oscillatory-integral form in H^1 and H^2 norms with dimension-independent rates (Theorem 3.6), and that Sobolev functions can be approximated by standard RaNNs with uniformly distributed hidden weights at rates that overcome the curse of dimensionality under sufficient regularity (Theorem 3.9). The numerical part reports PIELM experiments for the heat equation, the Black-Scholes model, and the Heston model in dimensions up to 100, with reported relative errors of a few percent and computation times of seconds to a few minutes.
Significance. If the main theorems were correct, the paper would make a useful contribution: dimension-independent H^1/H^2 approximation rates for randomized neural networks would supply a formal justification for PIELM-type methods in high-dimensional PDEs, complementing existing L^∞ results for random features. The proof strategy is constructive and the numerical experiments are extensive, which are strengths. However, the central proof contains a false equality in Eq. (3.12), and the constant I_ℓ in Theorem 3.6 is not well defined as written. These are load-bearing defects: Theorems 3.6 and 3.9 are not proven in the submitted manuscript, so the paper's main theoretical claim is currently unsupported.
major comments (3)
- [Section 3.1, Eq. (3.12)] The equality asserted in Eq. (3.12) is false for finite ε. Writing F(a)=∫_0^a H_ε(y)dy, the second integral equals F(x·ξ+u)−F(x·ξ+u−B(ξ)), so the difference between the two expressions is −F(x·ξ+u−B(ξ)), which is nonzero because H_ε is strictly positive. The error estimates in (3.21)–(3.41) are derived for the first integral, while the Monte-Carlo network in (3.47) implements the second integral. The proof therefore misses an unquantified boundary-layer contribution to u−u_ε and, consequently, to the networks in Theorems 3.6 and 3.9. The authors need either to replace the false identity by a valid bound for the boundary-layer difference in H^1(H^2), or to modify the definition of u_ε and H_ε so that the two forms are genuinely equal.
- [Theorem 3.6, Eq. (3.44)] The constant I_ℓ is not well defined as stated. Under Assumption 3.3, F(r)=2∫_0^{-r}πB(s)^{-1}ds is negative for r>0, and F(1)−F(−1) is always negative for any positive density πB. Hence the bracket [1+F(M‖ξ‖_1)+F(1)−F(−1)] can be negative, making the integrand in (3.44) negative and I_ℓ imaginary or undefined. This affects the validity of the error bound in Theorem 3.6. The definition should be corrected, for example by using a positive interval integral such as ∫_{-r}^0 πB(s)^{-1}ds, or by taking absolute values, and the proof of (3.50) should be checked with the corrected constant.
- [Theorem 3.9, Eq. (3.59) and Eq. (3.76)] The H^1 convergence rate stated in Theorem 3.9 is inconsistent with the proof. The theorem statement gives an exponent (s−d/2)/(6s+9), while the proof in Eq. (3.76) and Remark 3.10 both give (s−d/2)/(4s+6). These differ by a factor of 3/2 in the denominator. Since the H^1 rate is one of the two main advertised results, this is not a cosmetic typo in a lemma; the theorem statement must be corrected to match the derivation, or the derivation must be redone if (3.59) is the intended claim.
minor comments (6)
- [Eq. (3.22)] In the display following Eq. (3.22), the notation G(ε) appears in the integrand where G(ξ) is meant; the same confusion appears in the preceding line's bound.
- [Section 4, Tables 1–5] All numerical results appear to be single runs with no standard deviations or multiple random seeds. Given that the hidden weights A and B are random, reporting statistics over independent trials would substantially strengthen the empirical claims.
- [Section 4.1, Eq. (4.4)] The notation W·σ(A·(x,t)+b) is informal for the sum over features; the dimension of the bias vector b and the role of W₀ in Eq. (3.55) should be clarified.
- [Section 4.2] The quantity denoted E_r^G in Tables 3 and 4 is not defined in the text; it appears to be the relative L² error, but this should be stated explicitly.
- [Remark 3.10] The remark interprets the rates as tending to 1/4 and 1/10, but the abstract and introduction do not mention that these limits are obtained only as s→∞; a sentence clarifying the regularity requirement for dimension-independent rates would prevent overstatement.
- [Introduction] The sentence 'there are no available results on the approximation error of using randomized neural networks in high-dimensional PDEs' is too strong in view of the cited work of Gonon [11] on Black-Scholes type PDEs; the novelty should be phrased as Sobolev-norm approximation for RaNNs specifically.
Circularity Check
No significant circularity; theoretical rates rest on externally cited representation formulas and explicit error bounds, not on fitted parameters or self-citation chains.
full rationale
The derivation chain in Section 3 is not circular. The starting representation formula (3.8) is taken from Gonon (2023) and the Monte Carlo concentration lemma (Lemma 3.5) from Grohs et al. (2023), both external to the present authors; these are independent supports, and the paper's assumptions (Assumptions 3.1-3.3, 3.7) do not include the target approximation rates. The constants I_l and I*_l are upper bounds derived from the Fourier representation, not fitted to data, and the network weights W are constructed measurably rather than optimized against the target u. The Sobolev rates in Theorems 3.6 and 3.9 are proved by chaining the representation formula, the smooth-Heaviside approximation error, and a Monte Carlo estimate; no step redefines the target as the model. The numerical experiments in Section 4 are benchmarks against known exact or reference solutions and are not used to derive the theoretical rates. Self-citations (e.g., [4]-[6]) appear only as background on PINN training and error analysis and are not load-bearing. The contested equality (3.12) and the I_l sign issue are mathematical gaps or potential errors in the proof as written, not circular reductions: the second integral form is asserted, not shown to be the conclusion by construction. Even if the proof is incomplete, the argument is not equivalent to its inputs. Therefore no circularity is present.
Assumptions & free parameters
free parameters (4)
- gamma (Theorem 3.6 and 3.9 rate parameter) =
gamma in (0,1/2), e.g., tending to 1/2
- epsilon and delta (smoothing and boundary-layer parameters) =
powers of N and R chosen in the proofs
- beta_1, beta_2 (boundary scaling hyperparameters) =
5, 10, 50, 100 depending on dimension and problem
- Hidden weight sampling range =
U(-0.01, 0.01) for heat, U(-0.1, 0.1) for finance
assumptions (4)
- standard math Representation formula (3.8): u(x) = ∫∫ (x·ξ+u)_+ α(ξ,u) du dξ for Fourier-type functions (Gonon 2023, eq. 10).
- domain assumption Assumptions 3.1, 3.2, 3.3: A and B have strictly positive Lebesgue densities, Y uniform on [0,2M||A||_1+1], and F(r) finite.
- standard math Fourier inversion for u ∈ H^{4+s}(R^d): u(x)=∫ e^{ix·ξ} G(ξ)dξ with G related to the Fourier transform of u.
- ad hoc to paper Equation (3.12): ∫_0^{x·ξ+u} H_ε(y)dy = ∫_0^{B(ξ)} H_ε(x·ξ+u−y)dy for finite ε.
Cite this review
Pith. "Pith review of Approximation Theory and Applications of Randomized Neural Networks for Solving High-Dimensional PDEs." pith.science (2026). https://pith.science/paper/ECGJB44O
@misc{pith2026250112145,
author = {Pith},
title = {Pith review of: Approximation Theory and Applications of Randomized Neural Networks for Solving High-Dimensional PDEs},
year = {2026},
howpublished = {\url{https://pith.science/paper/ECGJB44O}},
note = {Machine review of arXiv:2501.12145}
}
abstract
We present approximation results and numerical experiments for the use of randomized neural networks within physics-informed extreme learning machines to efficiently solve high-dimensional PDEs, demonstrating both high accuracy and low computational cost. Specifically, we prove that RaNNs can approximate certain classes of functions, including Sobolev functions, in the $H^2$-norm at dimension-independent convergence rates, thereby alleviating the curse of dimensionality. Numerical experiments are provided for the high-dimensional heat equation, the Black-Scholes model, and the Heston model, demonstrating the accuracy and efficiency of randomized neural networks.
Reference graph
Works this paper leans on
-
[1]
A. R. Barron. Universal approximation bounds for superp ositions of a sigmoidal function. IEEE Transactions on Information theory , 39(3):930–945, 1993
work page 1993
-
[2]
J. Chen, X. Chi, Z. Yang, et al. Bridging traditional and m achine learning-based algorithms for solving PDEs: the random feature method. J Mach Learn , 1:268–98, 2022
work page 2022
-
[3]
H. Dang and F. W ang. Local randomized neural networks wit h hybridized discontinuous Petrov–Galerkin methods for Stokes–Darcy flows. Physics of Fluids , 36(8), 2024
work page 2024
-
[4]
T. De Ryck, F. Bonnet, S. Mishra, and E. de Bézenac. An oper ator preconditioning perspective on training in physics-informed machine learning. International Conference on Learning Representations , 2024
work page 2024
-
[5]
T. De Ryck and S. Mishra. Error analysis for physics infor med neural networks (PINNs) approximating Kol- mogorov PDEs. Advances in Computational Mathematics , 48(6):79, 2022
work page 2022
-
[6]
T. De Ryck and S. Mishra. Generic bounds on the approximat ion error for physics-informed (and) operator learning. Advances in Neural Information Processing Systems 35 , 2022
work page 2022
-
[7]
S. Dong and Z. Li. Local extreme learning machines and dom ain decomposition for solving linear and nonlinear partial differential equations. Computer Methods in Applied Mechanics and Engineering , 387:114129, 2021
work page 2021
-
[8]
V. Dwivedi and B. Srinivasan. Physics informed extreme l earning machine (pielm)–a rapid method for the numerical solution of partial differential equations. Neurocomputing, 391:96–118, 2020. RANDOMIZED NEURAL NETWORKS FOR HIGH-DIMENSIONAL PDES 19
work page 2020
Show all 43 references
-
[9]
W. E, J. Han, and A. Jentzen. Deep learning-based numeric al methods for high-dimensional parabolic partial differential equations and backward stochastic differentia l equations. Communications in Mathematics and Statistics, 5(4):349–380, 2017
2017
-
[10]
Gallicchio and S
C. Gallicchio and S. Scardapane. Deep randomized neura l networks. In Recent Trends in Learning From Data: Tutorials from the INNS Big Data and Deep Learning Conference (INNSBDDL2019), pages 43–68. Springer, 2020
2020
-
[11]
L. Gonon. Random feature neural networks learn black-s choles type pdes without curse of dimensionality. Journal of Machine Learning Research , 24(189):1–51, 2023
2023
-
[12]
Gonon, L
L. Gonon, L. Grigoryeva, and J.-P. Ortega. Approximati on bounds for random neural networks and reservoir systems. The Annals of Applied Probability , 33(1):28–69, 2023
2023
-
[13]
Grohs, F
P. Grohs, F. Hornung, A. Jentzen, and P. Von W urstemberg er. A proof that artificial neural networks overcome the curse of dimensionality in the numerical approximation of Black–Scholes partial differential equations , volume 284. American Mathematical Society, 2023
2023
-
[14]
Huang, Q.-Y
G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew. Extreme learnin g machine: a new learning scheme of feedforward neural networks. In 2004 IEEE international joint conference on neural network s (IEEE Cat. No. 04CH37541) , volume 2, pages 985–990. Ieee, 2004
2004
-
[15]
Huang, Q.-Y
G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew. Extreme learnin g machine: theory and applications. Neurocomputing, 70(1-3):489–501, 2006
2006
-
[16]
Hutzenthaler, A
M. Hutzenthaler, A. Jentzen, T. Kruse, and T. A. Nguyen. A proof that rectified deep neural networks overcome the curse of dimensionality in the numerical approximation of semilinear heat equations. SN partial differential equations and applications , 1(2):1–34, 2020
2020
-
[17]
Jentzen, D
A. Jentzen, D. Salimova, and T. W elti. A proof that deep a rtificial neural networks overcome the curse of dimensionality in the numerical approximation of Kolmogor ov partial differential equations with constant diffusion and nonlinear drift coefficients. ArXiv preprint arXiv:180...
2018 arXiv
-
[18]
Krishnapriyan, A
A. Krishnapriyan, A. Gholami, S. Zhe, R. Kirby, and M. W. Mahoney. Characterizing possible failure modes in physics-informed neural networks. Advances in Neural Information Processing Systems , 34:26548–26560, 2021
2021
-
[19]
Kutyniok, P
G. Kutyniok, P. Petersen, M. Raslan, and R. Schneider. A theoretical analysis of deep neural networks and parametric PDEs. Constructive Approximation, pages 1–53, 2021
2021
-
[20]
I. E. Lagaris, A. Likas, and P. G. D. Neural-network meth ods for boundary value problems with irregular boundaries. IEEE Transactions on Neural Networks , 11:1041–1049, 2000
2000
-
[21]
I. E. Lagaris, A. Likas, and D. I. Fotiadis. Artificial ne ural networks for solving ordinary and partial differential equations. IEEE Transactions on Neural Networks , 9(5):987–1000, 2000
2000
-
[22]
K. O. Lye, S. Mishra, and D. Ray. Deep learning observabl es in computational fluid dynamics. Journal of Computational Physics , page 109339, 2020
2020
-
[23]
K. O. Lye, S. Mishra, D. Ray, and P. Chandrashekar. Itera tive surrogate model optimization (ISMO): An active learning algorithm for pde constrained optimizatio n with deep neural networks. Computer Methods in Applied Mechanics and Engineering , 374:113575, 2021
2021
-
[24]
Mishra and R
S. Mishra and R. Molinaro. Physics informed neural netw orks for simulating radiative transfer. Journal of Quantitative Spectroscopy and Radiative Transfer , 270:107705, 2021
2021
-
[25]
Mishra and R
S. Mishra and R. Molinaro. Estimates on the generalizat ion error of physics-informed neural networks for approximating a class of inverse problems for pdes. IMA Journal of Numerical Analysis , 42:981–1022, 2022
2022
-
[26]
Mishra and R
S. Mishra and R. Molinaro. Estimates on the generalizat ion error of physics informed neural networks (PINNs) for approximating PDEs. IMA Journal of Numerical Analysis , 2022
2022
-
[27]
N. H. Nelsen and A. M. Stuart. The random feature model fo r input-output maps between banach spaces. SIAM Journal on Scientific Computing , 43(5):A3212–A3243, 2021
2021
-
[28]
Pao and Y
Y.-H. Pao and Y. Takefuji. Functional-link net computi ng: theory, system architecture, and functionalities. Computer, 25(5):76–79, 1992
1992
-
[29]
Rahimi and B
A. Rahimi and B. Recht. Random features for large-scale kernel machines. Advances in neural information processing systems, 20, 2007
2007
-
[30]
Rahimi and B
A. Rahimi and B. Recht. Uniform approximation of functi ons with random bases. In 2008 46th annual allerton conference on communication, control, and computing , pages 555–561. IEEE, 2008
2008
-
[31]
Rahimi and B
A. Rahimi and B. Recht. W eighted sums of random kitchen s inks: Replacing minimization with randomization in learning. Advances in neural information processing systems , 21, 2008
2008
-
[32]
Raissi and G
M. Raissi and G. E. Karniadakis. Hidden physics models: Machine learning of nonlinear partial differential equations. Journal of Computational Physics , 357:125–141, 2018
2018
-
[33]
Raissi, P
M. Raissi, P. Perdikaris, and G. E. Karniadakis. Physic s-informed neural networks: A deep learning frame- work for solving forward and inverse problems involving non linear partial differential equations. Journal of Computational Physics , 378:686–707, 2019
2019
-
[34]
Schiassi, R
E. Schiassi, R. Furfaro, C. Leake, M. De Florio, H. Johns ton, and D. Mortari. Extreme theory of functional connections: A fast physics-informed neural network metho d for solving ordinary and partial differential equations. Neurocomputing, 457:334–356, 2021
2021
-
[35]
Schwab and J
C. Schwab and J. Zech. Deep learning in high dimension: N eural network expression rates for generalized polynomial chaos expansions in uq. Analysis and Applications , 17(01):19–55, 2019
2019
-
[36]
Shang and F
Y. Shang and F. W ang. Randomized neural networks with Pe trov–Galerkin methods for solving linear elasticity and Navier–Stokes equations. Journal of Engineering Mechanics , 150(4):04024010, 2024. 20 RANDOMIZED NEURAL NETWORKS FOR HIGH-DIMENSIONAL PDES
2024
-
[37]
Shang, F
Y. Shang, F. W ang, and J. Sun. Randomized neural network with Petrov-Galerkin methods for solving linear and nonlinear partial differential equations. Communications in Nonlinear Science and Numerical Simulati on, 127:107518, 2023
2023
-
[38]
J. Sun, S. Dong, and F. W ang. Local randomized neural net works with discontinuous Galerkin methods for partial differential equations. Journal of Computational and Applied Mathematics , 445:115830, 2024
2024
-
[39]
Sun and F
J. Sun and F. W ang. Local randomized neural networks wit h discontinuous Galerkin methods for diffusive- viscous wave equation. Computers and Mathematics with Applications , 154:128–137, 2024
2024
-
[40]
Y. Sun, A. Gilbert, and A. Tewari. On the approximation p roperties of random ReLU features. ArXiv preprint arXiv:1810.04374, 2018
2018 arXiv
-
[41]
R. Tanios. Physics informed neural networks in computa tional finance: High dimensional forward and inverse option pricing. Master’s thesis, ETH Zurich, 2021
2021
-
[42]
W ang, Y
S. W ang, Y. Teng, and P. Perdikaris. Understanding and m itigating gradient flow pathologies in physics- informed neural networks. SIAM Journal on Scientific Computing , 43(5):A3055–A3081, 2021
2021
-
[43]
W ang, X
S. W ang, X. Yu, and P. Perdikaris. When and why PINNs fail to train: A neural tangent kernel perspective. Journal of Computational Physics , 449:110768, 2022
2022
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.