REVIEW 4 major objections 5 minor 125 references
LazyDINO: Fast, scalable, and efficiently amortized Bayesian inversion via structure-exploiting and surrogate-driven measure transport
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Under 1,000 offline solves now beat Laplace at Bayesian inversion.
desk verdict A genuinely useful combination of derivative-informed surrogates and lazy maps with convincing numerics, but the theory as written does not cover the implemented objective. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the lazy map $T_\theta=(I-P)+D_r\,\mathcal{T}_\theta E_r$, which leaves the prior untouched in the complement of a $d_r$-dimensional subspace and transports the whitened latent coordinates by $\mathcal{T}_\theta$. It is driven by a DIPNet ridge-function surrogate $V g_w(E_r\cdot)$, whose parameter encoder comes from the eigenproblem $H_A\psi_j=\lambda_j\psi_j$ and whose weights are trained with the derivative-informed objective that includes the latent Jacobian $J_r^{(j)}=V^*D G(m^{(j)})D_r$. The machinery transfers the expensive likelihood evaluation into a cheap neural evaluation in $\mathbb{R}^{d_r}$, and the theory ties the resulting posterior error to the eigenvalue tail sum and the Sobolev error of the surrogate.
What would settle it
Run LazyDINO on an inverse problem with a deliberately slow eigenvalue decay so that $\sum_{j>d_r}\lambda_j$ is not small, and check whether the posterior mean/covariance errors and the ANIS effective sample size degrade as predicted by Theorem 3.1; a second check is to swap Remark 2's zero-mean objective for the true conditional-expectation ridge function and see whether the reported advantage reverses at small sample counts.
Extended reading notes
Core claim
The central claim is that surrogate-driven lazy-map variational inference becomes both accurate and amortizable when the surrogate is co-designed with the latent structure of the posterior update. Concretely, the PtO map is replaced by a ridge function $G(m)\approx V g_w(E_r m)$ with $E_r$ projecting onto the leading $d_r$ eigenfunctions of the prior-preconditioned Gauss-Newton Hessian $H_A=\mathbb{E}_\mu[D_H G^* D_H G]$, and $g_w$ is trained with joint samples of the map and its Jacobian. Theorem 3.1 bounds the expected forward KL error by the eigenvalue tail sum plus a latent-representation error, and Theorem 3.2 with Corollary 3.3 show that the derivative-informed $H^1_\mu$ loss controls the gradient error and optimality gap of the lazy-map objective. The numerical sections show these bounds are tight enough in practice: with $d_r=200$ and fewer than 1,000 offline samples, LazyDINO outperforms the Laplace posterior in moment, density, and sampling diagnostics on two infinite-dimensional PDE inverse problems.
Load-bearing premise
The whole construction assumes the data move the posterior almost entirely within a 200-dimensional derivative-informed subspace, so the discarded eigenvalue tail of the prior-preconditioned Gauss-Newton Hessian is negligible; if informative directions fall outside it, the ridge surrogate cannot see them and the posterior error bound degrades.
Editorial extensions
If this is right
- The offline surrogate cost is amortized across every future data set sharing the same PtO map and prior, since the online phase only optimizes the cheap latent transport map.
- Posterior sampling and density evaluation become as fast as evaluating the trained lazy map, which enables real-time uncertainty quantification for digital twins and experimental design.
- Controlling the surrogate Jacobian, not just the map values, is what makes surrogate-driven transport optimization reliable; standard $L^2_\mu$ training can fail at 16,000 samples where derivative-informed training succeeds at 1,000.
- Because the parameter dimension is reduced to $d_r=200$ before any neural network or transport map is trained, the method's offline and online costs are independent of the discretization dimension of the PDE parameter field.
Reading between the lines
- The same derivative-informed ridge-function architecture should accelerate other query-intensive algorithms that differentiate through the PtO map, such as Bayesian optimal experimental design and PDE-constrained optimization under uncertainty, where the paper's optimality-gap bound would carry over.
- If the eigenvalue tail decays fast, the theory suggests an adaptive strategy: increase $d_r$ until the tail sum in (33) falls below the target posterior error, making LazyDINO's guarantee quantitative rather than heuristic.
- Remark 2's use of the zero prior mean instead of the conditional expectation leaves a testable gap: at very small sample budgets the practical objective differs from the theoretically optimal ridge function, and the paper's empirical choice suggests a bias-variance trade-off worth isolating in controlled experiments.
- A natural stress test is to push the method to multiple independent observations per parameter, where the posterior concentrates and the derivative-informed subspace may need to grow; the eigenvalue-tail criterion predicts exactly when the lazy-map ansatz breaks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces LazyDINO, a two-phase amortized Bayesian inversion method. In an offline phase, a derivative-informed reduced-basis neural operator (RB-DINO) surrogate of the parameter-to-observable map is trained using joint samples of the map and its Jacobian. In an online phase, this surrogate drives the training of a lazy map, a transport map whose nonlinearity acts only on a low-dimensional derivative-informed latent subspace. The authors claim two main theoretical results: (i) the DIPNet architecture and derivative-informed training minimize upper bounds on the expected surrogate posterior approximation error and on the expected optimality gap of surrogate-driven lazy-map optimization (Theorems 3.1, 3.2 and Corollary 3.3), and (ii) numerically, LazyDINO achieves high posterior accuracy at substantially lower offline cost than LazyNO, SBAI, LazyMap, and the Laplace approximation. The numerical study covers two nonlinear PDE-constrained inverse problems with four observation instances each and a wide range of posterior diagnostics.
Significance. The paper addresses an important practical problem: amortized Bayesian inversion for expensive PDE-governed models, where posterior approximation must be cheap after the model evaluations are performed offline. The co-design of the reduced basis, the neural surrogate, and the lazy-map transport is conceptually appealing, and the numerical comparison is unusually thorough: it includes ground-truth MCMC, moment discrepancies, density-based diagnostics, marginal visualizations, and careful computational-cost accounting. The algorithmic pseudocode and appendices are also detailed and would allow replication. If the theoretical characterization were fully established, the paper would make a solid contribution to surrogate-driven variational inference. However, as it stands, a load-bearing gap in the equivalence between the full-space and latent rKL objectives, together with an acknowledged deviation between the implemented and analyzed objectives, means that the central theoretical claims are not established for the algorithm as implemented.
major comments (4)
- [Section 2.5, Proposition 2.1] The equivalence in Eq. (22) is not valid under the stated assumptions. The full-space objective contains E_{z∼π,m⊥∼µ⊥}[1/2||V^*G(D_r Tθ(z)+m⊥)-V^*y||^2], while the latent objective contains E_{z∼π}[1/2||gopt(Tθ(z))-V^*y||^2]. The difference is E_{z∼π}[1/2 Var_{m⊥}(V^*G(D_r Tθ(z)+m⊥))], which depends on θ unless the conditional covariance of V^*G given the reduced coordinate is independent of that coordinate. No such condition is stated or verified. Since Theorem 3.2 and Corollary 3.3 operate in the latent formulation, this gap is load-bearing for the paper's main theoretical claims.
- [Remark 2 and Section 3.5, Eq. (43)] The implemented LazyDINO objective sets the complementary-space sample m⊥ to E[µ⊥]=0 rather than integrating over µ⊥. The surrogate-driven objective in Eq. (43a) is therefore E_z[1/2||g_w(Tθ(z))-V^*y||^2 + regularizer], not the expectation of the potential appearing in Eq. (25). This is a different variational problem from the one analyzed in Corollary 3.3. The text acknowledges the deviation but supplies neither a condition under which the two objectives coincide nor numerical evidence that the resulting bias is small. Consequently, claim (C1), that derivative-informed training minimizes the expected optimality gap of surrogate-driven lazy-map optimization, is not established for the algorithm as implemented.
- [Corollary 3.3] The optimality-gap result depends on assumptions (i) and (ii): the surrogate minimizer eθ^{y,†} must lie in a ball around the true minimizer θ^{y,†}, and the true objective must be locally strongly convex on that ball, γ-a.e. These are assumptions about problem- and surrogate-dependent quantities, and no verification is provided. The surrounding text states that derivative-informed learning minimizes the expected optimality gap without repeating these caveats, so the corollary should be presented as a conditional statement rather than an unconditional justification of the method.
- [Section 6.2, Figures 8 and 10] The numerical comparison is extensive, but the role of the chosen reduced basis dimension dr=200 and the MC approximation of the derivative-informed subspace with only 1000 samples is not investigated. The theoretical bounds in Theorem 3.1 and Eq. (33) assume the exact eigenbasis of the expected Gauss-Newton Hessian, while the experiments use a fixed sample-based approximation. This is not fatal, but a sensitivity study or at least a discussion of the approximation gap would be needed to connect the theory quantitatively to the reported results.
minor comments (5)
- [Abstract and Section 7] The abstract claims LazyDINO "consistently outperforms Laplace approximation" with fewer than 1000 samples, but Figure 8 shows that for Example I the Laplace baseline achieves lower covariance error than all methods, and Section 7 itself notes this exception. The abstract should be qualified.
- [Section 4, contribution list (C2)] There is a typo: "Scabalility" should be "Scalability".
- [Table 3] The note in Table 3 says "LazyDINO (1k) achieves smaller relative mean error than LazyDINO (128k)", but the comparison is with LazyMap (128k); the label should be corrected.
- [Section 6.1, Figures 6 and 7] The figures report a "statistical anomaly" in RB-DINO training at 500 samples. Since this anomaly propagates to the posterior comparisons, the paper should state whether the anomaly reflects a single seed or a systematic effect, and ideally report uncertainty over training seeds.
- [Section 5.4] The transport-map training schedule is reported as a list of (iterations, batch size, learning rate) tuples, but the notation is dense; a short table or clearer formatting would improve readability.
Circularity Check
No significant circularity: the theoretical bounds are genuine inequalities and the training losses coincide with the bound terms by design, which is a consistency result rather than a circular derivation.
full rationale
The derivation chain is not circular. Theorem 3.1 (Section 3.1) proves a genuine inequality, E_y DKL(mu^y || emu^y) <= Tr((I-P) H_A (I-P)) + E_z ||gopt - g||^2. Choosing the derivative-informed eigenbasis to minimize the first term and training the neural network to minimize the second term matches the bound's terms, but the bound itself does not presuppose DIPNet or the training loss; it is derived from H^1_mu regularity and the conditional-expectation optimality of eGopt. Similarly, Theorem 3.2 and Corollary 3.3 prove that the expected rKL gradient error and the expected optimality gap are controlled by the pi-weighted Sobolev norm of (gopt - g, grad gopt - grad g); the derivative-informed objective (40) is an unbiased single-sample estimator of exactly that norm, so minimizing it targets the bound. This is a consistency result, not circularity: the theorems provide the inequalities, and the losses are chosen to match the right-hand sides. The numerical superiority claim is benchmarked against MCMC ground truth and the Laplace baseline, so it is not forced by construction. The paper does cite prior work by overlapping authors ([24], [29], [30]), but those are used as proven mathematical results and algorithmic building blocks, and the core inequalities are re-proved in Appendices D.1-D.2; no load-bearing argument reduces to an unverified self-citation. Two correctness risks should be noted separately: Proposition 2.1's equality of full-space and latent rKL objectives requires the conditional covariance of G given the reduced coordinates to be independent of the reduced coordinate, a condition not stated in the paper; and Remark 2 explicitly replaces the m_perp expectation by E[mu_perp] = 0, so the implemented objective differs from the analyzed one. Neither is a circular step; both are gaps between assumptions and implementation. Verdict: no significant circularity.
Assumptions & free parameters
free parameters (4)
- reduced basis dimension dr =
200
- neural surrogate width and depth =
7 layers, 400 units
- lazy map IAF size =
30 layers, 400 hidden units
- stochastic optimization schedule =
Adamax batches 200-7500, learning rates 5e-3 to 5e-4; Adam with 1500 epochs
assumptions (6)
- standard math Gaussian prior distribution (Assumption 2.1)
- standard math Additive Gaussian noise (Assumption 2.2)
- domain assumption H1_mu-differentiable PtO map (Assumption 2.3)
- domain assumption Subspace concentration: data uninformative in Ker(P)
- domain assumption Local strong convexity and boundedness assumptions in Corollary 3.3
- ad hoc to paper Implementation uses E[mu_perp]=0 instead of the optimal conditional expectation (Remark 2)
Cite this review
Pith. "Pith review of LazyDINO: Fast, scalable, and efficiently amortized Bayesian inversion via structure-exploiting and surrogate-driven measure transport." pith.science (2026). https://pith.science/paper/76AIRJZO
@misc{pith2026241112726,
author = {Pith},
title = {Pith review of: LazyDINO: Fast, scalable, and efficiently amortized Bayesian inversion via structure-exploiting and surrogate-driven measure transport},
year = {2026},
howpublished = {\url{https://pith.science/paper/76AIRJZO}},
note = {Machine review of arXiv:2411.12726}
}
read the original abstract
We present LazyDINO, a transport map variational inference method for fast, scalable, and efficiently amortized solutions of high-dimensional nonlinear Bayesian inverse problems with expensive parameter-to-observable (PtO) maps. Our method consists of an offline phase in which we construct a derivative-informed neural surrogate of the PtO map using joint samples of the PtO map and its Jacobian. During the online phase, when given observational data, we seek rapid posterior approximation using surrogate-driven training of a lazy map [Brennan et al., NeurIPS, (2020)], i.e., a structure-exploiting transport map with low-dimensional nonlinearity. The trained lazy map then produces approximate posterior samples or density evaluations. Our surrogate construction is optimized for amortized Bayesian inversion using lazy map variational inference. We show that (i) the derivative-based reduced basis architecture [O'Leary-Roseberry et al., Comput. Methods Appl. Mech. Eng., 388 (2022)] minimizes the upper bound on the expected error in surrogate posterior approximation, and (ii) the derivative-informed training formulation [O'Leary-Roseberry et al., J. Comput. Phys., 496 (2024)] minimizes the expected error due to surrogate-driven transport map optimization. Our numerical results demonstrate that LazyDINO is highly efficient in cost amortization for Bayesian inversion. We observe one to two orders of magnitude reduction of offline cost for accurate posterior approximation, compared to simulation-based amortized inference via conditional transport and conventional surrogate-driven transport. In particular, LazyDINO outperforms Laplace approximation consistently using fewer than 1000 offline samples, while other amortized inference methods struggle and sometimes fail at 16,000 offline samples.
Figures
Figures from the paper (22 more)
Reference graph
Works this paper leans on
-
[1]
A. M. Stuart, Inverse problems: A Bayesian perspective, Acta Numerica 19 (2010) 451–559
2010
-
[2]
Bui-Thanh, O
T. Bui-Thanh, O. Ghattas, J. Martin, G. Stadler, A computational framework for infinite-dimensional bayesian inverse problems. part i: The linearized case, with application to global seismic inversion, 2013
2013
-
[3]
Petra, J
N. Petra, J. Martin, G. Stadler, O. Ghattas, A computational framework for infinite-dimensional bayesian inverse problems, part ii: Stochastic newton mcmc with application to ice sheet flow inverse problems, SIAM Journal on Scientific Computing 36 (2014) A1525–A1555
2014
-
[4]
Ghattas, K
O. Ghattas, K. Willcox, Learning physics-based models from data: perspectives from inverse problems and model reduction, Acta Numerica 30 (2021) 445–554
2021
-
[5]
M. G. Kapteyn, J. V. Pretorius, K. E. Willcox, A probabilistic graphical model foundation for enabling predictive digital twins at scale, Nature Computational Science 1 (2021) 337–347
2021
-
[6]
X. Huan, J. Jagalur, Y. Marzouk, Optimal experimental design: Formulations and computations, Acta Numerica 33 (2024) 715–840
2024
-
[7]
Baptista, Y
R. Baptista, Y. Marzouk, O. Zahm, On the representation and learning of monotone triangular transport maps, Foun- dations of Computational Mathematics (2023)
2023
-
[8]
Q. Liu, D. Wang, Stein variational gradient descent: A general purpose bayesian inference algorithm, 2019
2019
Show all 125 references
-
[9]
P. Chen, K. Wu, J. Chen, T. O’Leary-Roseberry, O. Ghattas, Projected stein variational newton: A fast and scalable bayesian inference method in high dimensions, Advances in Neural Information Processing Systems 32 (2019)
2019
-
[10]
Detommaso, T
G. Detommaso, T. Cui, A. Spantini, Y. Marzouk, R. Scheichl, A Stein variational Newton method, 2018
2018
-
[11]
D. J. Rezende, S. Mohamed, Variational inference with normalizing flows, arXiv:1505.05770 (2015)
2015 arXiv
-
[12]
Papamakarios, E
G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mohamed, B. Lakshminarayanan, Normalizing flows for probabilistic modeling and inference, arXiv preprint arXiv:1912.02762 (2019)
2019 arXiv
-
[13]
M. C. Brennan, D. Bigoni, O. Zahm, A. Spantini, Y. Marzouk, Greedy inference with structure-exploiting lazy maps, 2020
2020
-
[14]
T. Cui, J. Martin, Y. M. Marzouk, A. Solonen, A. Spantini, Likelihood-informed dimension reduction for nonlinear inverse problems, Inverse Problems 30 (2014) 114015
2014
-
[15]
P. G. Constantine, Active Subspaces, Society for Industrial and Applied Mathematics, Philadelphia, PA, 2015
2015
-
[16]
Bui-Thanh, O
T. Bui-Thanh, O. Ghattas, Analysis of the Hessian for inverse scattering problems. Part I: Inverse shape scattering of acoustic waves, Inverse Problems 28 (2012) 055001
2012
-
[17]
Bui-Thanh, O
T. Bui-Thanh, O. Ghattas, Analysis of the Hessian for inverse scattering problems. Part II: Inverse medium scattering of acoustic waves, Inverse Problems 28 (2012) 055002
2012
-
[18]
Bui-Thanh, O
T. Bui-Thanh, O. Ghattas, Analysis of the Hessian for inverse scattering problems. Part III: Inverse medium scattering of electromagnetic waves, Inverse Problems and Imaging 7 (2013) 1139–1155
2013
-
[19]
P. Chen, O. Ghattas, Hessian-based sampling for high-dimensional model reduction, International Journal for Uncertainty Quantification 9 (2019)
2019
-
[20]
P. Chen, U. Villa, O. Ghattas, Hessian-based adaptive sparse quadrature for infinite-dimensional bayesian inverse problems, Computer Methods in Applied Mechanics and Engineering 327 (2017) 147–172
2017
-
[21]
P. H. Flath, L. C. Wilcox, V. Ak¸ celik, J. Hill, B. van Bloemen Waanders, O. Ghattas, Fast algorithms for Bayesian uncertainty quantification in large-scale linear inverse problems based on low-rank partial Hessian approximations, SIAM Journal on Scientific Computing 33 (2011...
2011
-
[22]
Isaac, N
T. Isaac, N. Petra, G. Stadler, O. Ghattas, Scalable and efficient algorithms for the propagation of uncertainty from data through inference to prediction for large-scale problems, with application to flow of the Antarctic ice sheet, Journal of Computational Physics 296 (2015) 348–368
2015
-
[23]
Spantini, A
A. Spantini, A. Solonen, T. Cui, J. Martin, L. Tenorio, Y. Marzouk, Optimal low-rank approximations of bayesian linear inverse problems, SIAM Journal on Scientific Computing 37 (2015) A2451–A2487
2015
-
[24]
O’Leary-Roseberry, U
T. O’Leary-Roseberry, U. Villa, P. Chen, O. Ghattas, Derivative-informed projected neural networks for high-dimensional parametric maps governed by PDEs, Computer Methods in Applied Mechanics and Engineering 388 (2022) 114199
2022
-
[25]
Hesthaven, S
J. Hesthaven, S. Ubbiali, Non-intrusive reduced order modeling of nonlinear problems using neural networks, Journal of Computational Physics 363 (2018) 55–78
2018
-
[26]
Kovachki, Z
N. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. Stuart, A. Anandkumar, Neural operator: Learning maps between function spaces with applications to PDEs, Journal of Machine Learning Research 24 (2023) 1–97. 46
2023
-
[27]
N. B. Kovachki, S. Lanthaler, A. M. Stuart, Operator learning: Algorithms and analysis, arxiv.2402.15715 (2024)
2024 arXiv
-
[28]
D. Luo, T. O’Leary-Roseberry, P. Chen, O. Ghattas, Efficient PDE-constrained optimization under high-dimensional uncertainty using derivative-informed neural operators, arXiv preprint arXiv:2305.20053 (2023)
2023 arXiv
-
[29]
L. Cao, T. O’Leary-Roseberry, O. Ghattas, Derivative-informed neural operator acceleration of geometric MCMC for infinite-dimensional Bayesian inverse problems, arXiv preprint arXiv:2403.08220 (2024)
2024 arXiv
-
[30]
O’Leary-Roseberry, P
T. O’Leary-Roseberry, P. Chen, U. Villa, O. Ghattas, Derivative-informed neural operator: an efficient framework for high-dimensional parametric derivative learning, Journal of Computational Physics 496 (2024) 112555
2024
-
[31]
T. Cui, O. Zahm, Data-free likelihood-informed dimension reduction of Bayesian inverse problems, Inverse Problems 37 (2021) 045009
2021
-
[32]
O. Zahm, P. G. Constantine, C. Prieur, Y. M. Marzouk, Gradient-based dimension reduction of multivariate vector-valued functions, SIAM Journal on Scientific Computing 42 (2020) A534–A558
2020
-
[33]
P. S. Laplace, Memoir on the probability of the causes of events (1774). m´ emoires de math´ ematique et de physique, tome sixi` eme. (english translation by s. m. stigler), Statistical science 1 (1986) 364–378
1986
-
[34]
Bui-Thanh, C
T. Bui-Thanh, C. Burstedde, O. Ghattas, J. Martin, G. Stadler, L. C. Wilcox, Extreme-scale uq for bayesian inverse prob- lems governed by pdes, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’12, IEEE Compute...
2012
-
[35]
Y. M. Marzouk, H. N. Najm, Dimensionality reduction and polynomial chaos acceleration of bayesian inference in inverse problems, Journal of Computational Physics 228 (2009) 1862–1902
2009
-
[36]
O. Zahm, T. Cui, K. Law, A. Spantini, Y. Marzouk, Certified dimension reduction in nonlinear Bayesian inverse problems, Mathematics of Computation 91 (2022) 1789–1835
2022
-
[37]
Bigoni, Y
D. Bigoni, Y. Marzouk, C. Prieur, O. Zahm, Nonlinear dimension reduction for surrogate modeling using gradient information, Information and Inference: A Journal of the IMA 11 (2022) 1597–1639
2022
-
[38]
Y. M. Marzouk, H. N. Najm, L. A. Rahn, Stochastic spectral methods for efficient Bayesian solution of inverse problems, Journal of Computational Physics 224 (2007) 560–586
2007
-
[39]
Marzouk, D
Y. Marzouk, D. Xiu, A stochastic collocation approach to bayesian inference in inverse problems, COMMUNICATIONS IN COMPUTATIONAL PHYSICS 6 (2009) 826–847
2009
-
[40]
Farcas, J
I.-G. Farcas, J. Latz, E. Ullmann, T. Neckel, H.-J. Bungartz, Multilevel adaptive sparse Leja approximations for Bayesian inverse problems, SIAM Journal on Scientific Computing 42 (2020) A424–A451
2020
-
[41]
Galbally, K
D. Galbally, K. Fidkowski, K. Willcox, O. Ghattas, Non-linear model reduction for uncertainty quantification in large- scale inverse problems, International Journal for Numerical Methods in Engineering 81 (2010) 1581–1608
2010
-
[42]
Lieberman, K
C. Lieberman, K. Willcox, O. Ghattas, Parameter and state model reduction for large-scale statistical inverse problems, SIAM Journal on Scientific Computing 32 (2010) 2523–2542
2010
-
[43]
T. Cui, Y. M. Marzouk, K. E. Willcox, Data-driven model reduction for the Bayesian solution of inverse problems, International Journal for Numerical Methods in Engineering 102 (2015) 966–990
2015
-
[44]
Peherstorfer, K
B. Peherstorfer, K. Willcox, M. Gunzburger, Survey of multifidelity methods in uncertainty propagation, inference, and optimization, SIAM Review 60 (2018) 550–591
2018
-
[45]
M. B. Lykkegaard, T. J. Dodwell, C. Fox, G. Mingas, R. Scheichl, Multilevel delayed acceptance MCMC, SIAM/ASA Journal on Uncertainty Quantification 11 (2023) 1–30
2023
-
[46]
L. Cao, T. O’Leary-Roseberry, P. K. Jha, J. T. Oden, O. Ghattas, Residual-based error correction for neural operator accelerated infinite-dimensional Bayesian inverse problems, Journal of Computational Physics 486 (2023) 112104
2023
-
[47]
Bhattacharya, B
K. Bhattacharya, B. Hosseini, N. B. Kovachki, A. M. Stuart, Model reduction and neural network for parametric PDEs, The SMAI Journal of computational mathematics 7 (2021) 121–157
2021
-
[48]
Fresca, A
S. Fresca, A. Manzoni, POD-DL-ROM: Enhancing deep learning-based reduced order models for nonlinear parametrized PDEs by proper orthogonal decomposition, Computer Methods in Applied Mechanics and Engineering 388 (2022) 114181
2022
-
[49]
O’Leary-Roseberry, X
T. O’Leary-Roseberry, X. Du, A. Chaudhuri, J. R. Martins, K. Willcox, O. Ghattas, Learning high-dimensional para- metric maps via reduced basis adaptive residual networks, Computer Methods in Applied Mechanics and Engineering 402 (2022) 115730
2022
-
[50]
L. Lu, P. Jin, G. Pang, Z. Zhang, G. E. Karniadakis, Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Nature Machine Intelligence 3 (2021) 218–229
2021
-
[51]
J. H. Seidman, G. Kissas, G. J. Pappas, P. Perdikaris, Variational autoencoding neural operators, arXiv preprint, arXiv.2302.10351 (2023)
2023 arXiv
-
[52]
Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anandkumar, Fourier neural operator for parametric partial differential equations, arXiv preprint, arXiv.2010.08895 (2021)
2021 arXiv
-
[53]
Q. Cao, S. Goswami, G. E. Karniadakis, LNO: Laplace neural operator for solving differential equations, arXiv preprint, arXiv.2303.10528 (2023)
2023 arXiv
-
[54]
Lanthaler, Z
S. Lanthaler, Z. Li, A. M. Stuart, The nonlocal neural operator: Universal approximation, arXiv preprint, arXiv.2304.13221 (2023)
2023 arXiv
-
[55]
Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anandkumar, Neural operator: Graph kernel network for partial differential equations, arXiv preprint, arXiv.2003.03485 (2020)
2020 arXiv
-
[56]
Z. Li, H. Zheng, N. Kovachki, D. Jin, H. Chen, B. Liu, K. Azizzadenesheli, A. Anandkumar, Physics-informed neural operator for learning partial differential equations, ACM / IMS Journal of Data Science (2024)
2024
-
[57]
S. Wang, H. Wang, P. Perdikaris, Learning the solution operator of parametric partial differential equations with physics- informed DeepONets, Science Advances 7 (2021) eabi8605
2021
-
[58]
J. Go, P. Chen, Accelerating Bayesian Optimal Experimental Design with Derivative-Informed Neural Operators, arXiv preprint arXiv:2312.14810 (2023). 47
2023 arXiv
-
[59]
J. Go, P. Chen, Sequential infinite-dimensional Bayesian optimal experimental design with derivative-informed latent attention neural operator, arXiv preprint arXiv:2409.09141 (2024)
2024 arXiv
-
[60]
Y. Qiu, N. Bridges, P. Chen, Derivative-enhanced deep operator network, arXiv.2402.19242 (2024)
2024 arXiv
-
[61]
E. G. Tabak, C. V. Turner, A family of nonparametric density estimation algorithms, Communications on Pure and Applied Mathematics 66 (2013) 145–164
2013
-
[62]
Kobyzev, S
I. Kobyzev, S. Prince, M. Brubaker, Normalizing flows: An introduction and review of current methods, IEEE Transac- tions on Pattern Analysis and Machine Intelligence (2020)
2020
-
[63]
L. Dinh, J. Sohl-Dickstein, S. Bengio, Density estimation using real NVP, arXiv:1605.08803 (2016)
2016 arXiv
-
[64]
D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, M. Welling, Improved variational inference with inverse autoregressive flow, in: D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, R. Garnett (Eds.), Advances in Neural Information Processing Systems 29, Curran As...
2016
-
[65]
Papamakarios, T
G. Papamakarios, T. Pavlakou, I. Murray, Masked autoregressive flow for density estimation, in: Advances in Neural Information Processing Systems, pp. 2338–2347
-
[66]
Huang, D
C.-W. Huang, D. Krueger, A. Lacoste, A. Courville, Neural autoregressive flows, in: International Conference on Machine Learning, PMLR, pp. 2078–2087
-
[67]
De Cao, I
N. De Cao, I. Titov, W. Aziz, Block neural autoregressive flow, arXiv preprint arXiv:1904.04676 (2019)
2019 arXiv
-
[68]
Jaini, K
P. Jaini, K. A. Selby, Y. Yu, Sum-of-squares polynomial flow, in: International Conference on Machine Learning, pp. 3009–3018
-
[69]
Daniels, M
H. Daniels, M. Velikova, Monotone and partially monotone neural networks, IEEE Transactions on Neural Networks 21 (2010) 906–917
2010
-
[70]
Wehenkel, G
A. Wehenkel, G. Louppe, Unconstrained monotonic neural networks, Advances in neural information processing systems 32 (2019)
2019
-
[71]
Knothe, Contributions to the theory of convex bodies., Michigan Mathematical Journal 4 (1957) 39–52
H. Knothe, Contributions to the theory of convex bodies., Michigan Mathematical Journal 4 (1957) 39–52
1957
-
[72]
Rosenblatt, Remarks on a Multivariate Transformation, The Annals of Mathematical Statistics 23 (1952) 470 – 472
M. Rosenblatt, Remarks on a Multivariate Transformation, The Annals of Mathematical Statistics 23 (1952) 470 – 472
1952
-
[73]
T. A. El Moselhy, Y. M. Marzouk, Bayesian inference with optimal maps, Journal of Computational Physics 231 (2012) 7815–7850
2012
-
[74]
Marzouk, T
Y. Marzouk, T. Moselhy, M. Parno, A. Spantini, Sampling via measure transport: An introduction, in: Handbook of Uncertainty Quantification, R. Ghanem, D. Higdon, and H. Owhadi, editors, Springer, 2016
2016
-
[75]
Spantini, D
A. Spantini, D. Bigoni, Y. Marzouk, Inference via low-dimensional couplings, The Journal of Machine Learning Research 19 (2018) 2639–2709
2018
-
[76]
Baptista, O
R. Baptista, O. Zahm, Y. Marzouk, An adaptive transport framework for joint and conditional density estimation, arXiv:2009.10303 (2020)
2020 arXiv
-
[77]
J. Zech, Y. Marzouk, Sparse approximation of triangular transports, Part II: The infinite-dimensional case, Constr. Approx. 55 (2022) 987–1036
2022
-
[78]
Westermann, J
J. Westermann, J. Zech, Measure transport via polynomial density surrogates, arXiv preprint (2023)
2023
-
[79]
J. Zech, Y. Marzouk, Sparse approximation of triangular transports, Part I: The finite-dimensional case, Constr. Approx. 55 (2022) 919–986
2022
-
[80]
Zeghal, F
J. Zeghal, F. Lanusse, A. Boucaud, B. Remy, E. Aubourg, Neural posterior estimation with differentiable simulators, 2022
2022
-
[81]
Brehmer, G
J. Brehmer, G. Louppe, J. Pavez, K. Cranmer, Mining gold from implicit models to improve likelihood-free inference, Proceedings of the National Academy of Sciences 117 (2020) 5242–5249
2020
-
[82]
Durkan, G
C. Durkan, G. Papamakarios, I. Murray, Sequential neural methods for likelihood-free inference, arXiv preprint arXiv:1811.08723 (2018)
2018 arXiv
-
[83]
Papamakarios, D
G. Papamakarios, D. Sterratt, I. Murray, Sequential neural likelihood: Fast likelihood-free inference with autoregressive flows, in: The 22nd International Conference on Artificial Intelligence and Statistics, PMLR, pp. 837–848
-
[84]
Greenberg, M
D. Greenberg, M. Nonnenmacher, J. Macke, Automatic posterior transformation for likelihood-free inference, in: Inter- national Conference on Machine Learning, PMLR, pp. 2404–2414
-
[85]
Papamakarios, I
G. Papamakarios, I. Murray, Fast ϵ-free inference of simulation models with bayesian conditional density estimation, 2018
2018
-
[86]
Baptista, L
R. Baptista, L. Cao, J. Chen, O. Ghattas, F. Li, Y. M. Marzouk, J. T. Oden, Bayesian model calibration for block copolymer self-assembly: Likelihood-free inference and expected information gain computation via measure transport, Journal of Computational Physics 503 (2024) 112844
2024
-
[87]
Ganguly, S
A. Ganguly, S. Jain, U. Watchareeruetai, Amortized variational inference: A systematic review, Journal of Artificial Intelligence Research 78 (2023) 167–215
2023
-
[88]
Soize, R
C. Soize, R. Ghanem, Physical systems with random uncertainties: Chaos representations with arbitrary probability measure, SIAM Journal on Scientific Computing 26 (2004) 395–410
2004
-
[89]
Coordinate transformation and polynomial chaos for the bayesian inference of a gaussian process with parametrized prior covariance function, Computer Methods in Applied Mechanics and Engineering 298 (2016) 205–228
2016
-
[90]
D. M. Blei, A. Kucukelbir, J. D. McAuliffe, Variational inference: A review for statisticians, Journal of the American statistical Association 112 (2017) 859–877
2017
-
[91]
D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, M. Welling, Improving variational inference with inverse autoregressive flow, 2017
2017
-
[92]
Bottou, F
L. Bottou, F. E. Curtis, J. Nocedal, Optimization methods for large-scale machine learning, SIAM review 60 (2018) 223–311
2018
-
[93]
Isaac, N
T. Isaac, N. Petra, G. Stadler, O. Ghattas, Scalable and efficient algorithms for the propagation of uncertainty from 48 data through inference to prediction for large-scale problems, with application to flow of the antarctic ice sheet, Journal of Computational Physics 296 (20...
2015
-
[94]
Villa, N
U. Villa, N. Petra, O. Ghattas, hIPPYlib: An extensible software framework for large-scale inverse problems governed by PDEs: Part I: Deterministic inversion and linearized Bayesian inference, ACM Transactions on Mathematical Software 47 (2021)
2021
-
[95]
Villa, T
U. Villa, T. O’Leary-Roseberry, A note on the relationship between pde-based precision operators and mat \’ern covari- ances, arXiv preprint arXiv:2407.00471 (2024)
2024 arXiv
-
[96]
Kirchhoff, D
J. Kirchhoff, D. Luo, T. O’Leary-Roseberry, O. Ghattas, Inference of Heterogeneous Material Properties via Infinite- Dimensional Integrated DIC, arXiv preprint arXiv:2408.10217 (2024)
2024 arXiv
-
[97]
D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, 2017
2017
-
[98]
Beskos, G
A. Beskos, G. Roberts, A. Stuart, J. Voss, MCMC methods for diffusion bridges, Stochastics and Dynamics 8 (2008) 319–350
2008
-
[99]
Agapiou, O
S. Agapiou, O. Papaspiliopoulos, D. Sanz-Alonso, A. M. Stuart, Importance sampling: Intrinsic dimension and compu- tational cost, 2017
2017
-
[100]
B. Li, T. Bengtsson, P. Bickel, Curse-of-dimensionality revisited: Collapse of importance sampling in very large scale systems (2005)
2005
-
[101]
F. M. Polo, R. Vicente, Effective sample size, dimensionality, and generalization in covariate shift adaptation, Neural Computing and Applications 35 (2020) 18187–18199
2020
-
[102]
Sanz-Alonso, Z
D. Sanz-Alonso, Z. Wang, Bayesian update with importance sampling: Required sample size, Entropy 23 (2020) 22
2020
-
[103]
Beskos, M
A. Beskos, M. Girolami, S. Lan, P. E. Farrell, A. M. Stuart, Geometric MCMC for infinite-dimensional inverse problems, Journal of Computational Physics 335 (2017) 327–351
2017
-
[104]
Nualart, The Malliavin calculus and related topics, volume 1995, Springer, 2006
D. Nualart, The Malliavin calculus and related topics, volume 1995, Springer, 2006
1995
-
[105]
V. I. Bogachev, Gaussian measures, 62, American Mathematical Soc., 1998
1998
-
[106]
Halko, P.-G
N. Halko, P.-G. Martinsson, J. A. Tropp, Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions, 2010
2010
-
[107]
Xiang, J
H. Xiang, J. Zou, Randomized algorithms for large-scale inverse problems with general regularizations, 2014
2014
-
[108]
A. K. Saibaba, J. Lee, P. K. Kitanidis, Randomized algorithms for generalized hermitian eigenvalue problems with application to computing karhunen-lo` eve expansion, 2015
2015
-
[109]
G. H. Golub, Q. Ye, An inverse free preconditioned krylov subspace method for symmetric generalized eigenvalue problems, SIAM Journal on Scientific Computing 24 (2002) 312–334
2002
-
[110]
Saad, Krylov subspace methods for solving large unsymmetric linear systems, Mathematics of Computation 37 (1981) 105–126
Y. Saad, Krylov subspace methods for solving large unsymmetric linear systems, Mathematics of Computation 37 (1981) 105–126
1981
-
[111]
D. C. Sorensen, Truncated qz methods for large scale generalized eigenvalue problems, Electron. Trans. Numer. Anal. 7 (1998) 141–162
1998
-
[112]
van den Eshof, G
J. van den Eshof, G. L. G. Sleijpen, Inexact krylov subspace methods for linear systems, SIAM Journal on Matrix Analysis and Applications 26 (2004) 125–153
2004
-
[113]
Chowdhury, The truncated lanczos algorithm for partial solution of the symmetric eigenproblem, Computers & Structures 6 (1976) 439–446
P. Chowdhury, The truncated lanczos algorithm for partial solution of the symmetric eigenproblem, Computers & Structures 6 (1976) 439–446
1976
-
[114]
Z. Yao, A. Gholami, S. Shen, M. Mustafa, K. Keutzer, M. Mahoney, Adahessian: An adaptive second order optimizer for machine learning, in: proceedings of the AAAI conference on artificial intelligence, volume 35, pp. 10665–10673
-
[115]
O’Leary-Roseberry, R
T. O’Leary-Roseberry, R. Bollapragada, Fast Unconstrained Optimization via Hessian Averaging and Adaptive Gradient Sampling Methods, arXiv preprint arXiv:2408.07268 (2024)
2024 arXiv
-
[116]
Bollapragada, R
R. Bollapragada, R. Byrd, J. Nocedal, Adaptive sampling strategies for stochastic optimization, SIAM Journal on Optimization 28 (2018) 3312–3343
2018
-
[117]
Newton, R
D. Newton, R. Bollapragada, R. Pasupathy, N. K. Yip, A retrospective approximation approach for smooth stochastic optimization, 2024
2024
-
[118]
Kretschmann, Are Minimizers of the Onsager–Machlup Functional Strong Posterior Modes?, SIAM/ASA Journal on Uncertainty Quantification 11 (2023) 1105–1138
R. Kretschmann, Are Minimizers of the Onsager–Machlup Functional Strong Posterior Modes?, SIAM/ASA Journal on Uncertainty Quantification 11 (2023) 1105–1138
2023
-
[119]
R. S. Dembo, S. C. Eisenstat, T. Steihaug, Inexact newton methods, SIAM Journal on Numerical Analysis 19 (1982) 400–408
1982
-
[120]
S. C. Eisenstat, H. F. Walker, Choosing the forcing terms in an inexact newton method, SIAM Journal on Scientific Computing 17 (1996) 16–32. Appendix A. Glossary of terminology amortized cost: Discounted cost, by spreading it out over the solution of additional problems. The m...
1996
-
[121]
, N ▷ Sample prior
m(j) ∼ µ, j = 1, . . . , N ▷ Sample prior
-
[122]
, N ▷ Evaluate PtO map/Jacobian
G(m(j)), DH G(m(j)), j = 1, . . . , N ▷ Evaluate PtO map/Jacobian
-
[123]
Create encoder/decoder: {ψk ∈ M }dr k=1 ← eigenvalue problem( DG(m(j)) NL i=1, Γ−1 n , C−1, ϵL or dr) ▷ (32), (E.1) Drz := Pdr k=1 zkψk, Er := D⊤ r C−1 ▷ (15), (17)
-
[124]
, N end Latent space: solving the generalized eigenvalue problem
Embed dataset: ▷ (38), (42) z(j) ← Erm(j), g(j) ← V ⊤Γ−1 n G(m(j)), J (j) r ← V ⊤Γ−1 n DG(m(j))Dr, j = 1, . . . , N end Latent space: solving the generalized eigenvalue problem. We describe here the computation of the eigenvalue problem in Algorithm 1 to find the reduced basis...
-
[125]
Appendix E.2
Train gw by minimizing an empirical risk: w∗ = argmin w∈RdW 1 N PN j=1 g(j) − gw z(j) 2 + J (j) r − ∇zgw z(j) 2 F| {z } include for H 1µ (RB-DINO) objective end Equipped with a sufficiently accurate neural network approximation to the optimal latent PtO map gopt, including acc...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.