Pith. sign in

REVIEW 4 major objections 5 minor 83 references

MCMC-Net: Accelerating Markov Chain Monte Carlo with Neural Networks for Inverse Problems

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read MCMC-Net replaces an expensive forward solver with a CNN surrogate and shows the resulting posterior converges to the true posterior while running tens of times faster.

desk verdict A substantial empirical study of CNN-surrogate MCMC on three imaging problems, but the 'similar accuracy' claim overreaches and the missing forward-error measurement keeps the theory from biting. read the letter →

arxiv 2412.16883 v2 pith:X4L3O7KT submitted 2024-12-22 math.NA cs.NA

classification math.NAcs.NA MSC 62G2035J1562F15
keywords BayesianinverseproblemsMarkovChainMonteCarloneuraloperatorsurrogateconvolutionalnetworkHellingerdistanceelectricalimpedancetomographydiffuseopticalquantitativephotoacoustic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes MCMC-Net, a method that replaces the expensive forward-model evaluation inside Markov Chain Monte Carlo with a trained convolutional neural network. The authors show that for three PDE-based imaging inverse problems—electrical impedance tomography, diffuse optical tomography, and quantitative photoacoustic tomography—the surrogate likelihood produces posterior reconstructions comparable to the classical finite-element likelihood while cutting computation time by roughly an order of magnitude, up to about 30x. The paper also proves two theoretical guarantees: the neural operator can approximate the forward map to arbitrary accuracy in an L2 sense, and as that approximation error goes to zero, the surrogate posterior converges to the true posterior in Hellinger distance. If correct, this provides a principled justification for using fast learned surrogates in Bayesian uncertainty quantification.

What carries the argument

The engine is the composition $G_\theta = A \circ E$: an encoder $E$ that samples the parameter field on a grid, and a CNN approximator $A$ that maps those samples to the measurement vector. Theorem 2 shows this composition can approximate the true forward map to any desired $L^2(\mu)$ accuracy, with the error split into encoding and approximation parts. Theorem 3 then converts that forward error into a Hellinger-distance bound between the true and surrogate posterior measures, using the Gaussian likelihood's exponential structure.

What would settle it

Run MCMC-FEM and MCMC-Net on the same held-out phantom, measure the CNN's forward error at the posterior samples, and compute the Hellinger distance between the two posterior distributions; the paper's Theorem 3 predicts this distance tends to zero as the measured forward error does, so a case where the forward error is small but the posteriors differ noticeably would show the theoretical mechanism is not what drives the numerical result.

Watch

Extended reading notes

Core claim

The central claim is that MCMC-Net's surrogate posterior, built from a CNN replacement $G_\theta$ of the true forward map $G$, converges to the true posterior in Hellinger distance when the $L^2(\mu)$-error $\|G-G_\theta\|$ tends to zero, and that in practice this convergence is close enough for reliable reconstructions. The paper derives this by splitting the error into encoding and approximation terms and coupling them to bounds on the likelihood potentials; numerically, MCMC-Net achieves mean absolute error and mean square error values on par with or better than MCMC-FEM for EIT and DOT, and slightly higher errors for QPAT, while reducing inversion times from 125 to 4 minutes for EIT, 34 to 3.6 minutes for DOT, and 10.7 to 6.5 minutes for QPAT.

Load-bearing premise

The argument stands on the trained network's forward error being small over the parameter regions the sampler actually visits, and on the observed data being bounded; the paper does not measure the network's held-out forward error and explicitly leaves training and generalization error bounds for future work.

Editorial extensions

If this is right

  • The surrogate posterior converges to the true posterior in Hellinger distance as the trained forward surrogate error goes to zero, so Bayesian credible intervals computed with MCMC-Net are asymptotically faithful.
  • For EIT, DOT, and QPAT, likelihood evaluations with the CNN take minutes instead of hours, making full MCMC-based uncertainty quantification practical for these imaging modalities.
  • The speedup grows with discretization dimension: MCMC-FEM inversion time rises exponentially with mesh resolution, while MCMC-Net rises only linearly.
  • Because the proof only relies on the forward error being small in $L^2(\mu)$, the method extends to any Bayesian inverse problem whose likelihood is expensive and whose forward map admits a continuous approximator.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the theory implies a practical quality check—measure the trained surrogate's forward error on samples drawn from the posterior or its high-probability region, since the Hellinger bound is controlled by that $L^2(\mu)$ error; the paper does not report this quantity.
  • Editorial inference: the slight accuracy loss seen for QPAT is consistent with a small but nonzero surrogate forward error; a testable extension would be to train with a loss weighted toward the posterior-relevant region and then re-measure the Hellinger distance.
  • Editorial inference: the same surrogate-posterior convergence argument should carry over to other MCMC variants that use the same likelihood ratio, such as Metropolis-adjusted Langevin or Hamiltonian Monte Carlo, provided the surrogate error is controlled on the regions those samplers explore.
  • Editorial inference: an implicit corollary is that reliability is bounded by distribution shift—if test-time parameter fields fall outside the training distribution, the surrogate error can spike and bias the posterior, a risk only partially addressed by the paper's out-of-distribution experiments.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes MCMC-Net, a method for accelerating MCMC-based Bayesian inverse problems by replacing the FEM forward solver with a CNN surrogate in the likelihood evaluation. The authors present a universal approximation theorem for the neural operator surrogate, a Hellinger-distance convergence result for the surrogate posterior, and numerical experiments on EIT, DOT, and QPAT comparing reconstruction accuracy and speed against a FEM-based MCMC baseline. The central empirical claim is that MCMC-Net performs similarly to the classical likelihood counterpart while providing significant speedups.

Significance. If fully supported, the paper would address a major practical bottleneck in Bayesian inverse problems—the repeated evaluation of expensive PDE-based forward models inside MCMC. The strengths are the clean theoretical setup, the breadth of numerical experiments across three imaging modalities, the sensitivity and robustness analyses, and the inclusion of out-of-distribution tests. However, the central empirical claim is contradicted by the paper's own tables in several scenarios, and the theoretical convergence guarantee is not connected to the trained networks because the surrogate forward error is never measured. The significance of the contribution is therefore currently limited.

major comments (4)
  1. [Sections 5.4 and 5.6, Tables 4 and 7] The abstract's claim that MCMC-Net 'performs similar' to the classical likelihood counterpart is contradicted by the paper's own results. In Table 4, the three-anomaly EIT case gives MAE 0.350951 and MSE 1.403706 for MCMC-Net versus 0.282551 and 1.130100 for MCMC-FEM. In Table 7, MCMC-FEM has lower MAE and MSE at every QPAT noise level (e.g., at 2% noise, MAE 0.001329 vs 0.001708 and MSE 0.000101 vs 0.000156), and the speedup is only about 1.6x (10.62 vs 6.61 minutes). Later text (Sections 5.8 and 6) even asserts that the neural-operator method 'outperforms' the traditional method in error metrics, which is the opposite of Table 7. These discrepancies affect the paper's central claim.
  2. [Section 4, Theorems 2–3; Section 5] The theoretical guarantee is an existence result: Theorem 2 states that for every ε there exists some encoder–approximator pair with ||G − Gθ||_{L2(µ)} ≤ ε. It does not certify that the specific CNNs trained in Section 5 have small forward error, and the paper never reports a held-out estimate of ||G − Gθ||_{L2(µ)} over the prior or over the posterior region sampled by the chain. Section 6 explicitly defers training and generalization error bounds to future work. Without this measurement, Theorem 3's Hellinger convergence cannot be invoked for the numerical results, and the QPAT results in Table 7 are consistent with a nontrivial surrogate bias. This is the main methodological gap.
  3. [Theorem 3, Eq. (44)] The proof of Theorem 3 uses the boundedness of y in Eq. (44) ('is bounded if y is bounded') without stating it as a hypothesis of the theorem. The same bound is used to control I2, so this is a load-bearing omission. The theorem statement should include the assumption that y is bounded, or the proof should be adjusted to avoid this requirement.
  4. [Theorem 2 proof, Eqs. (34)–(35)] The notation in the proof of Theorem 2 is imprecise: in Eq. (35) the expression ||G(Id − I_P)w|| should read ||G(w) − G(I_P w)||, and the integrals over the sets K and Z are written without explicitly specifying the restrictions of µ and the integrands. The argument can be repaired, but as written it is difficult to follow.
minor comments (5)
  1. [Section 4.1] There is a typo: 'the goal will; be' should be 'the goal will be'.
  2. [Theorem 2 statement] The statement 'there exists a finite-dimensional spaces' should be 'there exists a finite-dimensional space'.
  3. [Figure A.11 caption] The caption repeats 'the same number of layers (NL = 4)' for three successive triples of subfigures; the intended differences (e.g., number of neurons) are unclear and should be stated more precisely.
  4. [Section 5.8] The text asserts that MCMC-Net 'outperforms' MCMC-FEM in error metrics for QPAT, which is inconsistent with Table 7 and should be corrected.
  5. [Section 5.1] The acceptance probability is written as min(1, L(q*_Prop.) − L(q_i)) without showing the prior ratio; although pCN with a Gaussian prior can cancel prior terms, the expression as written is incomplete and should be clarified.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the surrogate-posterior convergence is a conditional continuity bound in the forward-model error, the CNN is trained on FEM forward data rather than on posterior targets, and the self-citations are background or baseline material rather than load-bearing premises.

full rationale

The derivation chain is not circular. The central theoretical claim, Theorem 3, shows that the Hellinger distance between the surrogate posterior mu_y^theta and the true posterior mu^y is controlled by ||G - G_theta||_{L2(mu)}: Eqs. (40)-(43) bound I1 and I2 by powers of this forward-operator error, so the posterior convergence is a continuity statement with respect to the surrogate's forward accuracy. The quantity being controlled is the same forward error that Theorem 2 asserts can be made small for some encoder/approximator pair; nothing in the theorem defines the posterior or the likelihood in terms of the surrogate posterior, and the proof cites external results (Marzouk and Xiu [79], Iglesias-Lu-Stuart [45], Lanthaler-Mishra-Karniadakis [74], Kovachki-Lanthaler-Mishra [75]). The numerical surrogates are trained on input-output pairs generated by FEM forward solves (Sec. 5.3), not on posterior samples or posterior summary statistics, so the inversion results are not fitted parameters renamed as predictions. The self-citations ([38], [46], [81], [82]) provide background, baselines, or pCN tuning heuristics and do not carry the load-bearing approximation or posterior-convergence argument. The paper's own limitation statements are relevant: Sec. 6 explicitly defers training and generalization error bounds to future work, and Theorem 3's proof uses boundedness of y in Eq. (44) without stating it as a hypothesis; moreover, the paper does not report a held-out ||G - G_theta||_{L2(mu)} measurement for the trained CNNs. These are genuine soundness or rigor gaps between the asymptotic theorem and the particular networks used in the experiments, but a missing measurement or an omitted hypothesis is not a circular reduction. No step of the derivation is equivalent to its inputs by construction.

Assumptions & free parameters 5 free parameters · 7 assumptions · 0 invented entities

The central claims rest on standard Bayesian inversion theory, on a compactness-and-approximation argument for the forward maps, and on an unverified premise that the trained CNN has small L2 error against the FEM forward model. The main hand-chosen quantities are the prior and network hyperparameters; the most fragile assumptions are the boundedness of y in Theorem 3 and the unmeasured surrogate error.

free parameters (5)
  • Matérn prior length scale ℓ = 0.4 (EIT), 0.2 (DOT)
    Chosen heuristically in Sections 5.1 and 5.3; controls the smoothness and regularization strength of the Gaussian process prior.
  • Matérn prior smoothness υ = 3
    Fixed for all problems; a modeling choice for the Gaussian process prior.
  • pCN proposal step size Δ = Adapted during burn-in to maintain about 25% acceptance
    Tuned per chain to approximate the optimal acceptance rate; affects mixing but not the target posterior.
  • CNN architecture and training hyperparameters = 4 layers, 16-64 neurons, minibatch 8-128, epochs 100-2000, learning rate 0.001
    Chosen by systematic experimentation (Section 5.3, Tables 6, A.8, A.9); the reported inversion results depend on these choices, though sensitivity analysis suggests mild dependence.
  • Level-set and star-shaped prior constants = σ_i, c_i for EIT; κ_i for QPAT
    Set by the problem and prior definitions; not fitted to data.
assumptions (7)
  • standard math Bayes' theorem for infinite-dimensional inverse problems (Theorem 1, from [40])
    Used to define the true posterior and the surrogate posterior in Eq. (3) and Eq. (7).
  • domain assumption The forward maps G for EIT, DOT, and QPAT are continuous and bounded on their parameter spaces
    Invoked in the proof of Theorem 2 via compactness and Lusin's theorem; standard for these level-set and star-shaped formulations, but not proved in this paper.
  • standard math Trigonometric interpolation error decays as ∥(Id−IP)w∥_{L∞} ≲ P^{−ξ(s)} for w in H^s
    Quoted from [75, Theorems 39 and 40] to control the encoder error T1 in Theorem 2.
  • standard math ReLU networks can uniformly approximate continuous functions on compact subsets of R^P
    Quoted from [77, Theorem 2] and [78, Theorem 1] to justify the approximator A.
  • standard math The pCN proposal is reversible with respect to the Gaussian prior, so only the likelihood ratio appears in the acceptance probability
    Used in Section 5.1 without proof; standard result for pCN.
  • ad hoc to paper The observed data y is bounded in Theorem 3
    The proof of Eq. (44) needs ∥y∥<∞, but the theorem statement and the QPAT white-noise model do not guarantee bounded data.
  • domain assumption Point evaluations plus trigonometric interpolation capture enough of the parameter field for the forward map
    The proof of Theorem 2 decomposes the error through DP∘EP; the implemented network has no decoder, so the analysis is conditional on the CNN approximating G∘DP on the encoded space.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MCMC-Net: Accelerating Markov Chain Monte Carlo with Neural Networks for Inverse Problems." pith.science (2026). https://pith.science/paper/X4L3O7KT

@misc{pith2026241216883,
  author       = {Pith},
  title        = {Pith review of: MCMC-Net: Accelerating Markov Chain Monte Carlo with Neural Networks for Inverse Problems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X4L3O7KT}},
  note         = {Machine review of arXiv:2412.16883}
}
read the original abstract

In many computational problems, using the Markov Chain Monte Carlo (MCMC) can be prohibitively time-consuming. We propose MCMC-Net, a simple yet efficient way to accelerate MCMC via neural networks. The key idea of our approach is to substitute the true likelihood function of the MCMC method with a neural operator based surrogate. We extensively evaluate the accuracy and speedup of our method on three different PDE-based inverse problems where likelihood computations are computationally expensive, namely electrical impedance tomography, diffuse optical tomography, and quantitative photoacoustic tomography. MCMC-Net performs similar to the classical likelihood counterpart but with a significant speedup. We conjecture that the method can be applied to any problem with a sufficiently expensive likelihood function. We also analyze MCMC-Net in a theoretical setting for the different use cases. We prove a universal approximation theorem-type result to show that the proposed network can approximate the mapping resulting from forward model evaluations to a desired accuracy. Furthermore, we establish convergence of the surrogate posterior to the true posterior under Hellinger distance.

Figures

Figures reproduced from arXiv: 2412.16883 by the authors.

Figure 1
Figure 1. Flowchart illustrating the MCMC-Net workflow [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. The true map G is approximated by a composition of two maps, encoder E and approximator A. Note that in the Bayesian inverse problem formulated in section 3 for EIT, the space Hs (Ω) := ¯ X is a separable Hilbert space, and H : X → Γ is a level set map from the space X to the space of piecewise constant admissible conductivities. In this case, we denote the forward problem as: y = GEIT(w) + η (30) where GEIT : X → R… view at source ↗
Figure 3
Figure 3. Reconstruction of electrical conductivity using MCMC-FEM and MCMC-Net for EIT (data ob [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: EIT: Convergence of the MCMC method for FEM and CNN. Data obtained at 1 % relative noise. [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]
Figure 5
Figure 5. Figure 5: Trace plots of the posterior samples of level-set functions using the proposed approach for EIT [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 6
Figure 6. Figure 6: Reconstruction of absorption coefficient ( [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: Reconstruction of the absorption coefficient using MCMC-FEM and MCMC-Net for QPAT. From [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]
Figure 8
Figure 8. Figure 8: This figure shows qualitative results corresponding to the quantitative results in Table 6, arranged [PITH_FULL_IMAGE:figures/full_fig_p027_8.png]
Figure 9
Figure 9. Figure 9: Comparison of inversion times between MCMC-FEM and MCMC-Net. For each method, 100,000 [PITH_FULL_IMAGE:figures/full_fig_p028_9.png]
Figure 8
Figure 8. Figure 8: We also report the inversion time using the MCMC-Net. [PITH_FULL_IMAGE:figures/full_fig_p041_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

83 extracted references · 57 canonical work pages

  1. [20]

    J. Li, Y. M. Marzouk, Adaptive construction of surrogates for the Bayesian solution of inverse problems, SIAM Journal on Scientific Computing 36 (3) (2014) A1163–A1186

  2. [23]

    Z. Gao, L. Yan, T. Zhou, Adaptive operator learning for infinite-dimensional Bayesian inverse problems, SIAM/ASA Journal on Uncertainty Quantification 12 (4) (2024) 1389– 1423

  3. [29]

    Deveney, E

    T. Deveney, E. Mueller, T. Shardlow, A deep surrogate approach to efficient Bayesian inversion in pde and integral equation models, arXiv preprint arXiv:1910.01547 (2019)

  4. [33]

    L. Cao, T. O’Leary-Roseberry, P. K. Jha, J. T. Oden, O. Ghattas, Residual-based error correction for neural operator accelerated infinite-dimensional Bayesian inverse problems, Journal of Computational Physics 486 (2023) 112104

  5. [1]

    A. E. Gelfand, A. F. Smith, Sampling-based approaches to calculating marginal densi- ties, Journal of the American statistical association 85 (410) (1990) 398–409

  6. [2]

    Metropolis, A

    N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller, E. Teller, Equation of state calculations by fast computing machines, The journal of chemical physics 21 (6) (1953) 1087–1092

  7. [3]

    W. K. Hastings, Monte carlo sampling methods using markov chains and their applica- tions (1970)

  8. [4]

    C. P. Robert, G. Casella, Monte carlo statistical methods (springer texts in statistics) (2005)

Show all 83 references
  1. [5]

    G. O. Roberts, J. S. Rosenthal, General state space markov chains and MCMC algo- rithms (2004)

  2. [6]

    Kaipio, E

    J. Kaipio, E. Somersalo, Statistical and computational inverse problems, Vol. 160, Springer Science & Business Media, 2006

  3. [7]

    Tarantola, Inverse problem theory and methods for model parameter estimation, SIAM, 2005

    A. Tarantola, Inverse problem theory and methods for model parameter estimation, SIAM, 2005

  4. [8]

    H. W. Engl, M. Hanke, A. Neubauer, Regularization of inverse problems, Vol. 375, Springer Science & Business Media, 1996

  5. [9]

    A. M. Stuart, Inverse problems: a Bayesian perspective, Acta numerica 19 (2010) 451– 559

  6. [10]

    Calvetti, E

    D. Calvetti, E. Somersalo, Inverse problems: From regularization to Bayesian in- ference, WIREs Computational Statistics 10 (3) (2018) e1427.arXiv:https:// wires.onlinelibrary.wiley.com/doi/pdf/10.1002/wics.1427,doi:https://doi. org/10.1002/wics.1427. URLhttps://wires.onlineli...

  7. [11]

    Cotter, G

    S. Cotter, G. Roberts, A. Stuart, D. White, MCMC methods for functions: Modifying old algorithms to make them faster, Statistical Science 28 (3) (2013) 424–446

  8. [12]

    Gelman, W

    A. Gelman, W. Gilks, G. Roberts, Weak convergence and optimal scaling of random walk metropolis algorithms, The annals of applied probability 7 (1) (1997) 110–120

  9. [13]

    Goodman, J

    J. Goodman, J. Weare, Ensemble samplers with affine invariance, Communications in applied mathematics and computational science 5 (1) (2010) 65–80

  10. [14]

    T. Cui, Y. M. Marzouk, K. E. Willcox, Data-driven model reduction for the Bayesian so- lution of inverse problems, International Journal for Numerical Methods in Engineering 102 (5) (2015) 966–990

  11. [15]

    T. Cui, Y. Marzouk, K. Willcox, Scalable posterior approximations for large-scale Bayesian inverse problems via likelihood-informed parameter and state reduction, Jour- nal of Computational Physics 315 (2016) 363–387. 29

  12. [16]

    Lieberman, K

    C. Lieberman, K. Willcox, O. Ghattas, Parameter and state model reduction for large- scale statistical inverse problems, SIAM Journal on Scientific Computing 32 (5) (2010) 2523–2542

  13. [17]

    Y. M. Marzouk, H. N. Najm, Dimensionality reduction and polynomial chaos accelera- tion of Bayesian inference in inverse problems, Journal of Computational Physics 228 (6) (2009) 1862–1902

  14. [18]

    Bui-Thanh, O

    T. Bui-Thanh, O. Ghattas, J. Martin, G. Stadler, A computational framework for infinite-dimensional Bayesian inverse problems part i: The linearized case, with appli- cation to global seismic inversion, SIAM Journal on Scientific Computing 35 (6) (2013) A2494–A2523

  15. [19]

    Schillings, B

    C. Schillings, B. Sprungk, P. Wacker, On the convergence of the laplace approximation and noise-level-robustness of laplace-based monte carlo methods for Bayesian inverse problems, Numerische Mathematik 145 (2020) 915–971

  16. [21]

    Yan, Y.-X

    L. Yan, Y.-X. Zhang, Convergence analysis of surrogate-based methods for Bayesian inverse problems, Inverse Problems 33 (12) (2017) 125001

  17. [22]

    L. Y. Zhou, et al., An adaptive surrogate modeling based on deep neural networks for large-scale Bayesian inverse problems, Communications in Computational Physics 28 (5) (2020) 2180–2205

  18. [24]

    J. Han, A. Jentzen, W. E, Solving high-dimensional partial differential equations using deep learning, Proceedings of the National Academy of Sciences 115 (34) (2018) 8505– 8510

  19. [25]

    Raissi, P

    M. Raissi, P. Perdikaris, G. E. Karniadakis, Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational physics 378 (2019) 686–707

  20. [26]

    Schwab, J

    C. Schwab, J. Zech, Deep learning in high dimension: Neural network expression rates for generalized polynomial chaos expansions in uq, Analysis and Applications 17 (01) (2019) 19–55

  21. [27]

    R. K. Tripathy, I. Bilionis, Deep uq: Learning deep neural network surrogate models for high dimensional uncertainty quantification, Journal of computational physics 375 (2018) 565–588

  22. [28]

    Y. Zhu, N. Zabaras, Bayesian deep convolutional encoder–decoder networks for surrogate modeling and uncertainty quantification, Journal of Computational Physics 366 (2018) 415–447. 30

  23. [30]

    L. Yan, T. Zhou, An acceleration strategy for randomize-then-optimize sampling via deep neural networks, Journal of Computational Mathematics 39 (6) (2021) 848–864

  24. [31]

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anandku- mar, Fourier neural operator for parametric partial differential equations, arXiv preprint arXiv:2010.08895 (2020)

  25. [32]

    L. Lu, P. Jin, G. Pang, Z. Zhang, G. E. Karniadakis, Learning nonlinear operators via deeponet based on the universal approximation theorem of operators, Nature machine intelligence 3 (3) (2021) 218–229

  26. [34]

    Genzel, J

    M. Genzel, J. Macdonald, M. M¨ arz, Solving inverse problems with deep neural networks– robustness included?, IEEE transactions on pattern analysis and machine intelligence 45 (1) (2022) 1119–1134

  27. [35]

    Raoni´ c, R

    B. Raoni´ c, R. Molinaro, T. De Ryck, T. Rohner, F. Bartolucci, R. Alaifari, S. Mishra, E. de B´ ezenac, Convolutional neural operators for robust and accurate learning of pdes, Advances in Neural Information Processing Systems 36 (2024)

  28. [36]

    Molinaro, Y

    R. Molinaro, Y. Yang, B. Engquist, S. Mishra, Neural inverse operators for solving pde inverse problems (2023).arXiv:2301.11167. URLhttps://arxiv.org/abs/2301.11167

  29. [37]

    S. U. Park, Estimation for compositional data using measurements from nonlinear sys- tems using artificial neural networks (2020).arXiv:2001.09040. URLhttps://arxiv.org/abs/2001.09040

  30. [38]

    Abhishek, T

    A. Abhishek, T. Strauss, A deeponet for inverting the neumann-to-dirichlet operator in electrical impedance tomography: An approximation theoretic perspective and numeri- cal results (2025).arXiv:2407.17182. URLhttps://arxiv.org/abs/2407.17182

  31. [39]

    Nickl, Bayesian Non-linear Statistical Inverse Problems, Zurich Lectures in Advanced Mathematics, EMS, 2023.doi:10.4171/ZLAM/30

    R. Nickl, Bayesian Non-linear Statistical Inverse Problems, Zurich Lectures in Advanced Mathematics, EMS, 2023.doi:10.4171/ZLAM/30. URLhttps://ems.press/books/zlam/260

  32. [40]

    Dashti, A

    M. Dashti, A. M. Stuart, The Bayesian approach to inverse problems, in: Handbook of uncertainty quantification. Vol. 1, 2, 3, Springer, Cham, 2017, pp. 311–428

  33. [41]

    Somersalo, M

    E. Somersalo, M. Cheney, D. Isaacson, Existence and uniqueness for electrode models for electric current computed tomography, SIAM J. Appl. Math. 52 (4) (1992) 1023–1040. doi:10.1137/0152060. URLhttps://doi.org/10.1137/0152060 31

  34. [42]

    M. M. Dunlop, A. M. Stuart, The Bayesian formulation of EIT: analysis and algorithms, Inverse Probl. Imaging 10 (4) (2016) 1007–1036.doi:10.3934/ipi.2016030. URLhttps://doi.org/10.3934/ipi.2016030

  35. [43]

    Cheney, D

    M. Cheney, D. Isaacson, J. C. Newell, Electrical impedance tomography, SIAM review 41 (1) (1999) 85–101

  36. [44]

    Borcea, Electrical impedance tomography, Inverse problems 18 (6) (2002) R99

    L. Borcea, Electrical impedance tomography, Inverse problems 18 (6) (2002) R99

  37. [45]

    M. A. Iglesias, Y. Lu, A. Stuart, A Bayesian level set method for geometric inverse problems, Interfaces Free Bound. 18 (2) (2016) 181–217.doi:10.4171/IFB/362. URLhttps://doi.org/10.4171/IFB/362

  38. [46]

    Abhishek, T

    A. Abhishek, T. Strauss, T. Khan, An optimal Bayesian estimator for absorption coef- ficient in diffuse optical tomography, SIAM Journal on Imaging Sciences 15 (2) (2022) 797–821

  39. [47]

    Natterer, F

    F. Natterer, F. W¨ ubbeling, Mathematical methods in image reconstruction, SIAM Monographs on Mathematical Modeling and Computation, Society for Industrial and Applied Mathematics (SIAM), Philadelphia, PA, 2001.doi:10.1137/1. 9780898718324. URLhttps://doi.org/10.1137/1.9780898718324

  40. [48]

    Harrach, On uniqueness in diffuse optical tomography, Inverse problems 25 (5) (2009) 055010

    B. Harrach, On uniqueness in diffuse optical tomography, Inverse problems 25 (5) (2009) 055010

  41. [49]

    B. M. Afkham, K. Knudsen, A. K. Rasmussen, T. Tarvainen, A Bayesian approach for consistent reconstruction of inclusions, Inverse Problems 40 (4) (2024) 045004

  42. [50]

    Abraham, R

    K. Abraham, R. Nickl, On statistical Calder´ on problems, Math. Stat. Learn. 2 (2) (2019) 165–216

  43. [51]

    Suhonen, A

    M. Suhonen, A. Pulkkinen, T. Tarvainen, Single-stage approach for estimating optical parameters in spectral quantitative photoacoustic tomography, JOSA A 41 (3) (2024) 527–542

  44. [52]

    Tarvainen, B

    T. Tarvainen, B. T. Cox, J. Kaipio, S. R. Arridge, Reconstructing absorption and scat- tering distributions in quantitative photoacoustic tomography, Inverse Problems 28 (8) (2012) 084009

  45. [53]

    Quarteroni, A

    A. Quarteroni, A. Valli, Numerical approximation of partial differential equations, Vol. 23, Springer Science & Business Media, 2008

  46. [54]

    G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, L. Yang, Physics- informed machine learning, Nature Reviews Physics 3 (6) (2021) 422–440

  47. [55]

    T. Chen, H. Chen, Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems, IEEE trans- actions on neural networks 6 (4) (1995) 911–917. 32

  48. [56]

    S. Cai, Z. Wang, L. Lu, T. A. Zaki, G. E. Karniadakis, Deepm&mnet: Inferring the elec- troconvection multiphysics fields based on operator approximation by neural networks, Journal of Computational Physics 436 (2021) 110296

  49. [57]

    Z. Mao, L. Lu, O. Marxen, T. A. Zaki, G. E. Karniadakis, Deepm&mnet for hyperson- ics: Predicting the coupled flow and finite-rate chemistry behind a normal shock using neural-network approximation of operators, Journal of computational physics 447 (2021) 110698

  50. [58]

    Bhattacharya, B

    K. Bhattacharya, B. Hosseini, N. B. Kovachki, A. M. Stuart, Model reduction and neural networks for parametric pdes, The SMAI journal of computational mathematics 7 (2021) 121–157

  51. [59]

    Kovachki, Z

    N. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. Stuart, A. Anand- kumar, Neural operator: Learning maps between function spaces with applications to pdes, Journal of Machine Learning Research 24 (89) (2023) 1–97

  52. [60]

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anand- kumar, Neural operator: Graph kernel network for partial differential equations, arXiv preprint arXiv:2003.03485 (2020)

  53. [61]

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, A. Stuart, K. Bhattacharya, A. Anand- kumar, Multipole graph neural operator for parametric partial differential equations, Advances in Neural Information Processing Systems 33 (2020) 6755–6766

  54. [62]

    Z. Li, H. Zheng, N. Kovachki, D. Jin, H. Chen, B. Liu, K. Azizzadenesheli, A. Anand- kumar, Physics-informed neural operator for learning partial differential equations, ACM/JMS Journal of Data Science 1 (3) (2024) 1–27

  55. [63]

    Pathak, S

    J. Pathak, S. Subramanian, P. Harrington, S. Raja, A. Chattopadhyay, M. Mardani, T. Kurth, D. Hall, Z. Li, K. Azizzadenesheli, et al., Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators, arXiv preprint arXiv:2202.11214 (2022)

  56. [64]

    Prasthofer, T

    M. Prasthofer, T. De Ryck, S. Mishra, Variable-input deep operator networks, arXiv preprint arXiv:2205.11404 (2022)

  57. [65]

    V. S. Fanaskov, I. V. Oseledets, Spectral neural operators, in: Doklady Mathematics, Vol. 108, Springer, 2023, pp. S226–S232

  58. [66]

    Kissas, J

    G. Kissas, J. H. Seidman, L. F. Guilhoto, V. M. Preciado, G. J. Pappas, P. Perdikaris, Learning operators with coupled attention, Journal of Machine Learning Research 23 (215) (2022) 1–63

  59. [67]

    Seidman, G

    J. Seidman, G. Kissas, P. Perdikaris, G. J. Pappas, Nomad: Nonlinear manifold decoders for operator learning, Advances in Neural Information Processing Systems 35 (2022) 5601–5613

  60. [68]

    M. V. de Hoop, M. Lassas, C. A. Wong, Deep learning architectures for nonlinear op- erator functions and nonlinear inverse problems, Mathematical Statistics and Learning 4 (1) (2022) 1–86. 33

  61. [69]

    M. V. de Hoop, N. B. Kovachki, N. H. Nelsen, A. M. Stuart, Convergence rates for learn- ing linear operators from noisy data, SIAM/ASA Journal on Uncertainty Quantification 11 (2) (2023) 480–513

  62. [70]

    Furuya, M

    T. Furuya, M. Puthawala, M. Lassas, M. V. de Hoop, Globally injective and bijective neural operators, Advances in Neural Information Processing Systems 36 (2024)

  63. [71]

    Cao, Choose a transformer: Fourier or galerkin, Advances in neural information processing systems 34 (2021) 24924–24940

    S. Cao, Choose a transformer: Fourier or galerkin, Advances in neural information processing systems 34 (2021) 24924–24940

  64. [72]

    Uslu Tuna, M

    H. Uslu Tuna, M. Sari, T. Cosgun, A discretization-free deep neural network-based approach for advection-dispersion-reaction mechanisms, Physica Scripta 99 (7) (2024) 076006.doi:10.1088/1402-4896/ad5258. URLhttps://dx.doi.org/10.1088/1402-4896/ad5258

  65. [73]

    H. U. Tuna, M. Sari, T. C. and, Unveiling advection–dominated interactions: Efficacy of neural networks in natural systems modelling, Numerical Heat Transfer, Part B: Fundamentals 0 (0) (2024) 1–15.arXiv:https://doi.org/10.1080/10407790.2024. 2392001,doi:10.1080/10407790.2024....

  66. [74]

    Lanthaler, S

    S. Lanthaler, S. Mishra, G. E. Karniadakis, Error estimates for deeponets: A deep learn- ing framework in infinite dimensions, Transactions of Mathematics and Its Applications 6 (1) (2022) tnac001

  67. [75]

    Kovachki, S

    N. Kovachki, S. Lanthaler, S. Mishra, On universal approximation and error bounds for fourier neural operators, Journal of Machine Learning Research 22 (290) (2021) 1–76

  68. [76]

    N. B. Kovachki, Machine Learning and Scientific Computing, California Institute of Technology, 2022

  69. [77]

    Yarotsky, Optimal approximation of continuous functions by very deep relu networks, in: S

    D. Yarotsky, Optimal approximation of continuous functions by very deep relu networks, in: S. Bubeck, V. Perchet, P. Rigollet (Eds.), Proceedings of the 31st Conference On Learning Theory, Vol. 75 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 639–649. URLhttps:/...

  70. [78]

    Zhou, Universality of deep convolutional neural networks, Applied and Compu- tational Harmonic Analysis 48 (2) (2020) 787–794.doi:https://doi.org/10.1016/j

    D.-X. Zhou, Universality of deep convolutional neural networks, Applied and Compu- tational Harmonic Analysis 48 (2) (2020) 787–794.doi:https://doi.org/10.1016/j. acha.2019.06.004. URLhttps://www.sciencedirect.com/science/article/pii/S1063520318302045

  71. [79]

    Marzouk, D

    Y. Marzouk, D. Xiu, A stochastic collocation approach to Bayesian inference in inverse problems, Commun. Comput. Phys. 6 (4) (2009) 826–847. URLhttps://global-sci.org/intro/article_detail/cicp/7708.html

  72. [80]

    Gelman, W

    A. Gelman, W. R. Gilks, G. O. Roberts, Weak convergence and optimal scaling of random walk metropolis algorithms, The Annals of Applied Probability 7(1) (1997)

  73. [81]

    Ahmad, T

    S. Ahmad, T. Strauss, S. Kupis, T. Khan, Comparison of statistical inversion with itera- tively regularized gauss newton method for image reconstruction in electrical impedance tomography, Applied Mathematics and Computation 358 (2019) 436–448. 34

  74. [82]

    Strauss, T

    T. Strauss, T. Khan, Statistical inversion in electrical impedance tomography using mixed total variation and non-convexℓ p regularization prior, J. Inverse Ill-Posed Probl. 23 (5) (2015) 529–542

  75. [83]

    Rasmussen, Inclusion-qPAT,https://github.com/akselkaastras/ inclusion-qPAT(2023)

    A. Rasmussen, Inclusion-qPAT,https://github.com/akselkaastras/ inclusion-qPAT(2023). 35 Appendix A. Sensitivity analysis of Network architectures for DOT and QPAT inversions. Earlier, we tabulated the results for how sensitive the Bayesian inversion in the EIT is to the chosen...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.