Pith. sign in

REVIEW 4 major objections 5 minor 125 references

LazyDINO: Fast, scalable, and efficiently amortized Bayesian inversion via structure-exploiting and surrogate-driven measure transport

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Under 1,000 offline solves now beat Laplace at Bayesian inversion.

desk verdict A genuinely useful combination of derivative-informed surrogates and lazy maps with convincing numerics, but the theory as written does not cover the implemented objective. read the letter →

arxiv 2411.12726 v1 pith:76AIRJZO submitted 2024-11-19 math.NA cs.LGcs.NAstat.COstat.ML

classification math.NAcs.LGcs.NAstat.COstat.ML MSC 65N2162F1565C6068T07
keywords Bayesianinverseproblemsvariationalinferencemeasuretransportlazymapsderivative-informedneuraloperatorsdimensionreductionamortizedPDE-constraineduncertaintyquantification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

LazyDINO attacks nonlinear Bayesian inverse problems whose parameter-to-observable (PtO) map is too expensive to evaluate inside a sampling loop. It builds one derivative-informed neural surrogate of the PtO map offline, then uses that surrogate to train a lazy map, a transport map that is nonlinear only in a low-dimensional latent subspace, for each new data set online. The paper proves that this two-step design is the right one for amortized inference: the reduced-basis architecture minimizes an upper bound on expected posterior error, and the derivative-informed training loss minimizes the expected optimality gap of the surrogate-driven transport optimization. In two PDE-governed examples, LazyDINO delivers posterior approximations that beat the Laplace approximation with fewer than 1,000 offline PDE solves, while conventional surrogate-driven and simulation-based amortized methods still struggle at 16,000 samples.

What carries the argument

The load-bearing object is the lazy map $T_\theta=(I-P)+D_r\,\mathcal{T}_\theta E_r$, which leaves the prior untouched in the complement of a $d_r$-dimensional subspace and transports the whitened latent coordinates by $\mathcal{T}_\theta$. It is driven by a DIPNet ridge-function surrogate $V g_w(E_r\cdot)$, whose parameter encoder comes from the eigenproblem $H_A\psi_j=\lambda_j\psi_j$ and whose weights are trained with the derivative-informed objective that includes the latent Jacobian $J_r^{(j)}=V^*D G(m^{(j)})D_r$. The machinery transfers the expensive likelihood evaluation into a cheap neural evaluation in $\mathbb{R}^{d_r}$, and the theory ties the resulting posterior error to the eigenvalue tail sum and the Sobolev error of the surrogate.

What would settle it

Run LazyDINO on an inverse problem with a deliberately slow eigenvalue decay so that $\sum_{j>d_r}\lambda_j$ is not small, and check whether the posterior mean/covariance errors and the ANIS effective sample size degrade as predicted by Theorem 3.1; a second check is to swap Remark 2's zero-mean objective for the true conditional-expectation ridge function and see whether the reported advantage reverses at small sample counts.

Watch

Extended reading notes

Core claim

The central claim is that surrogate-driven lazy-map variational inference becomes both accurate and amortizable when the surrogate is co-designed with the latent structure of the posterior update. Concretely, the PtO map is replaced by a ridge function $G(m)\approx V g_w(E_r m)$ with $E_r$ projecting onto the leading $d_r$ eigenfunctions of the prior-preconditioned Gauss-Newton Hessian $H_A=\mathbb{E}_\mu[D_H G^* D_H G]$, and $g_w$ is trained with joint samples of the map and its Jacobian. Theorem 3.1 bounds the expected forward KL error by the eigenvalue tail sum plus a latent-representation error, and Theorem 3.2 with Corollary 3.3 show that the derivative-informed $H^1_\mu$ loss controls the gradient error and optimality gap of the lazy-map objective. The numerical sections show these bounds are tight enough in practice: with $d_r=200$ and fewer than 1,000 offline samples, LazyDINO outperforms the Laplace posterior in moment, density, and sampling diagnostics on two infinite-dimensional PDE inverse problems.

Load-bearing premise

The whole construction assumes the data move the posterior almost entirely within a 200-dimensional derivative-informed subspace, so the discarded eigenvalue tail of the prior-preconditioned Gauss-Newton Hessian is negligible; if informative directions fall outside it, the ridge surrogate cannot see them and the posterior error bound degrades.

Editorial extensions

If this is right

  • The offline surrogate cost is amortized across every future data set sharing the same PtO map and prior, since the online phase only optimizes the cheap latent transport map.
  • Posterior sampling and density evaluation become as fast as evaluating the trained lazy map, which enables real-time uncertainty quantification for digital twins and experimental design.
  • Controlling the surrogate Jacobian, not just the map values, is what makes surrogate-driven transport optimization reliable; standard $L^2_\mu$ training can fail at 16,000 samples where derivative-informed training succeeds at 1,000.
  • Because the parameter dimension is reduced to $d_r=200$ before any neural network or transport map is trained, the method's offline and online costs are independent of the discretization dimension of the PDE parameter field.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same derivative-informed ridge-function architecture should accelerate other query-intensive algorithms that differentiate through the PtO map, such as Bayesian optimal experimental design and PDE-constrained optimization under uncertainty, where the paper's optimality-gap bound would carry over.
  • If the eigenvalue tail decays fast, the theory suggests an adaptive strategy: increase $d_r$ until the tail sum in (33) falls below the target posterior error, making LazyDINO's guarantee quantitative rather than heuristic.
  • Remark 2's use of the zero prior mean instead of the conditional expectation leaves a testable gap: at very small sample budgets the practical objective differs from the theoretically optimal ridge function, and the paper's empirical choice suggests a bias-variance trade-off worth isolating in controlled experiments.
  • A natural stress test is to push the method to multiple independent observations per parameter, where the posterior concentrates and the derivative-informed subspace may need to grow; the eigenvalue-tail criterion predicts exactly when the lazy-map ansatz breaks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces LazyDINO, a two-phase amortized Bayesian inversion method. In an offline phase, a derivative-informed reduced-basis neural operator (RB-DINO) surrogate of the parameter-to-observable map is trained using joint samples of the map and its Jacobian. In an online phase, this surrogate drives the training of a lazy map, a transport map whose nonlinearity acts only on a low-dimensional derivative-informed latent subspace. The authors claim two main theoretical results: (i) the DIPNet architecture and derivative-informed training minimize upper bounds on the expected surrogate posterior approximation error and on the expected optimality gap of surrogate-driven lazy-map optimization (Theorems 3.1, 3.2 and Corollary 3.3), and (ii) numerically, LazyDINO achieves high posterior accuracy at substantially lower offline cost than LazyNO, SBAI, LazyMap, and the Laplace approximation. The numerical study covers two nonlinear PDE-constrained inverse problems with four observation instances each and a wide range of posterior diagnostics.

Significance. The paper addresses an important practical problem: amortized Bayesian inversion for expensive PDE-governed models, where posterior approximation must be cheap after the model evaluations are performed offline. The co-design of the reduced basis, the neural surrogate, and the lazy-map transport is conceptually appealing, and the numerical comparison is unusually thorough: it includes ground-truth MCMC, moment discrepancies, density-based diagnostics, marginal visualizations, and careful computational-cost accounting. The algorithmic pseudocode and appendices are also detailed and would allow replication. If the theoretical characterization were fully established, the paper would make a solid contribution to surrogate-driven variational inference. However, as it stands, a load-bearing gap in the equivalence between the full-space and latent rKL objectives, together with an acknowledged deviation between the implemented and analyzed objectives, means that the central theoretical claims are not established for the algorithm as implemented.

major comments (4)
  1. [Section 2.5, Proposition 2.1] The equivalence in Eq. (22) is not valid under the stated assumptions. The full-space objective contains E_{z∼π,m⊥∼µ⊥}[1/2||V^*G(D_r Tθ(z)+m⊥)-V^*y||^2], while the latent objective contains E_{z∼π}[1/2||gopt(Tθ(z))-V^*y||^2]. The difference is E_{z∼π}[1/2 Var_{m⊥}(V^*G(D_r Tθ(z)+m⊥))], which depends on θ unless the conditional covariance of V^*G given the reduced coordinate is independent of that coordinate. No such condition is stated or verified. Since Theorem 3.2 and Corollary 3.3 operate in the latent formulation, this gap is load-bearing for the paper's main theoretical claims.
  2. [Remark 2 and Section 3.5, Eq. (43)] The implemented LazyDINO objective sets the complementary-space sample m⊥ to E[µ⊥]=0 rather than integrating over µ⊥. The surrogate-driven objective in Eq. (43a) is therefore E_z[1/2||g_w(Tθ(z))-V^*y||^2 + regularizer], not the expectation of the potential appearing in Eq. (25). This is a different variational problem from the one analyzed in Corollary 3.3. The text acknowledges the deviation but supplies neither a condition under which the two objectives coincide nor numerical evidence that the resulting bias is small. Consequently, claim (C1), that derivative-informed training minimizes the expected optimality gap of surrogate-driven lazy-map optimization, is not established for the algorithm as implemented.
  3. [Corollary 3.3] The optimality-gap result depends on assumptions (i) and (ii): the surrogate minimizer eθ^{y,†} must lie in a ball around the true minimizer θ^{y,†}, and the true objective must be locally strongly convex on that ball, γ-a.e. These are assumptions about problem- and surrogate-dependent quantities, and no verification is provided. The surrounding text states that derivative-informed learning minimizes the expected optimality gap without repeating these caveats, so the corollary should be presented as a conditional statement rather than an unconditional justification of the method.
  4. [Section 6.2, Figures 8 and 10] The numerical comparison is extensive, but the role of the chosen reduced basis dimension dr=200 and the MC approximation of the derivative-informed subspace with only 1000 samples is not investigated. The theoretical bounds in Theorem 3.1 and Eq. (33) assume the exact eigenbasis of the expected Gauss-Newton Hessian, while the experiments use a fixed sample-based approximation. This is not fatal, but a sensitivity study or at least a discussion of the approximation gap would be needed to connect the theory quantitatively to the reported results.
minor comments (5)
  1. [Abstract and Section 7] The abstract claims LazyDINO "consistently outperforms Laplace approximation" with fewer than 1000 samples, but Figure 8 shows that for Example I the Laplace baseline achieves lower covariance error than all methods, and Section 7 itself notes this exception. The abstract should be qualified.
  2. [Section 4, contribution list (C2)] There is a typo: "Scabalility" should be "Scalability".
  3. [Table 3] The note in Table 3 says "LazyDINO (1k) achieves smaller relative mean error than LazyDINO (128k)", but the comparison is with LazyMap (128k); the label should be corrected.
  4. [Section 6.1, Figures 6 and 7] The figures report a "statistical anomaly" in RB-DINO training at 500 samples. Since this anomaly propagates to the posterior comparisons, the paper should state whether the anomaly reflects a single seed or a systematic effect, and ideally report uncertainty over training seeds.
  5. [Section 5.4] The transport-map training schedule is reported as a list of (iterations, batch size, learning rate) tuples, but the notation is dense; a short table or clearer formatting would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the theoretical bounds are genuine inequalities and the training losses coincide with the bound terms by design, which is a consistency result rather than a circular derivation.

full rationale

The derivation chain is not circular. Theorem 3.1 (Section 3.1) proves a genuine inequality, E_y DKL(mu^y || emu^y) <= Tr((I-P) H_A (I-P)) + E_z ||gopt - g||^2. Choosing the derivative-informed eigenbasis to minimize the first term and training the neural network to minimize the second term matches the bound's terms, but the bound itself does not presuppose DIPNet or the training loss; it is derived from H^1_mu regularity and the conditional-expectation optimality of eGopt. Similarly, Theorem 3.2 and Corollary 3.3 prove that the expected rKL gradient error and the expected optimality gap are controlled by the pi-weighted Sobolev norm of (gopt - g, grad gopt - grad g); the derivative-informed objective (40) is an unbiased single-sample estimator of exactly that norm, so minimizing it targets the bound. This is a consistency result, not circularity: the theorems provide the inequalities, and the losses are chosen to match the right-hand sides. The numerical superiority claim is benchmarked against MCMC ground truth and the Laplace baseline, so it is not forced by construction. The paper does cite prior work by overlapping authors ([24], [29], [30]), but those are used as proven mathematical results and algorithmic building blocks, and the core inequalities are re-proved in Appendices D.1-D.2; no load-bearing argument reduces to an unverified self-citation. Two correctness risks should be noted separately: Proposition 2.1's equality of full-space and latent rKL objectives requires the conditional covariance of G given the reduced coordinates to be independent of the reduced coordinate, a condition not stated in the paper; and Remark 2 explicitly replaces the m_perp expectation by E[mu_perp] = 0, so the implemented objective differs from the analyzed one. Neither is a circular step; both are gaps between assumptions and implementation. Verdict: no significant circularity.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

The central claim depends on the choice of the dimension-reducing subspace (dr=200) and on the assumption that the posterior update is confined to that subspace. The bounds are proven under Gaussian prior/noise and Sobolev differentiability assumptions. No invented entities are introduced.

free parameters (4)
  • reduced basis dimension dr = 200
    Selected by hand for both examples and fixed across all compared methods. The eigenvalue tail at dr=200 is implicitly assumed small, directly controlling the expressiveness of the ridge function surrogate.
  • neural surrogate width and depth = 7 layers, 400 units
    Capacity of the MLP gw affects the representation error and the Sobolev norm in the bounds. Chosen by common practice, not by a principled criterion.
  • lazy map IAF size = 30 layers, 400 hidden units
    Flexibility of the transport map; affects how well the latent posterior can be approximated. Set to the same architecture for all compared methods.
  • stochastic optimization schedule = Adamax batches 200-7500, learning rates 5e-3 to 5e-4; Adam with 1500 epochs
    The training schedule for the transport map and surrogate influences the achieved optimality gap and generalization. Chosen empirically and reported in Section 5.
assumptions (6)
  • standard math Gaussian prior distribution (Assumption 2.1)
    The method assumes a Gaussian prior on the parameter. Non-Gaussian priors require a change of variables as noted in Remark 1.
  • standard math Additive Gaussian noise (Assumption 2.2)
    The likelihood is Gaussian, leading to the quadratic potential in Eq. (5).
  • domain assumption H1_mu-differentiable PtO map (Assumption 2.3)
    The Malliavin derivative and the Hessian-based bounds require the PtO map to be differentiable in the Gaussian Sobolev sense. This is a technical assumption that holds for typical PDE maps under sufficient regularity.
  • domain assumption Subspace concentration: data uninformative in Ker(P)
    The ridge function approximation (11) and the use of the derivative-informed subspace assume the posterior-to-prior update is confined to a dr-dimensional subspace; the eigenvalue tail in (33) must be small. This is the main structural assumption of the method.
  • domain assumption Local strong convexity and boundedness assumptions in Corollary 3.3
    The optimality gap bound assumes the rKL objective is locally strongly convex near the true optimum and the surrogate minimizer lands within that region. These conditions are not verified for the numerical examples.
  • ad hoc to paper Implementation uses E[mu_perp]=0 instead of the optimal conditional expectation (Remark 2)
    The numerical method replaces the theoretical objective based on the optimal ridge function with a single-sample unconditional prior mean, changing the optimization problem that the theorems bound. This is an acknowledged deviation, not an assumption in the published theory.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LazyDINO: Fast, scalable, and efficiently amortized Bayesian inversion via structure-exploiting and surrogate-driven measure transport." pith.science (2026). https://pith.science/paper/76AIRJZO

@misc{pith2026241112726,
  author       = {Pith},
  title        = {Pith review of: LazyDINO: Fast, scalable, and efficiently amortized Bayesian inversion via structure-exploiting and surrogate-driven measure transport},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/76AIRJZO}},
  note         = {Machine review of arXiv:2411.12726}
}
read the original abstract

We present LazyDINO, a transport map variational inference method for fast, scalable, and efficiently amortized solutions of high-dimensional nonlinear Bayesian inverse problems with expensive parameter-to-observable (PtO) maps. Our method consists of an offline phase in which we construct a derivative-informed neural surrogate of the PtO map using joint samples of the PtO map and its Jacobian. During the online phase, when given observational data, we seek rapid posterior approximation using surrogate-driven training of a lazy map [Brennan et al., NeurIPS, (2020)], i.e., a structure-exploiting transport map with low-dimensional nonlinearity. The trained lazy map then produces approximate posterior samples or density evaluations. Our surrogate construction is optimized for amortized Bayesian inversion using lazy map variational inference. We show that (i) the derivative-based reduced basis architecture [O'Leary-Roseberry et al., Comput. Methods Appl. Mech. Eng., 388 (2022)] minimizes the upper bound on the expected error in surrogate posterior approximation, and (ii) the derivative-informed training formulation [O'Leary-Roseberry et al., J. Comput. Phys., 496 (2024)] minimizes the expected error due to surrogate-driven transport map optimization. Our numerical results demonstrate that LazyDINO is highly efficient in cost amortization for Bayesian inversion. We observe one to two orders of magnitude reduction of offline cost for accurate posterior approximation, compared to simulation-based amortized inference via conditional transport and conventional surrogate-driven transport. In particular, LazyDINO outperforms Laplace approximation consistently using fewer than 1000 offline samples, while other amortized inference methods struggle and sometimes fail at 16,000 offline samples.

Figures

Figures reproduced from arXiv: 2411.12726 by the authors.

Figure 1
Figure 1. Overview of the RB-DINO construction. Step 2: RB-DINO Surrogate-Driven Lazy Map Optimization Given Observational Data y 0.0 0.2 0.4 0.6 0.8 1.0 Observed data y ∈ R dy gw : R dr → R dy V ∗ 1 2 ∥(gw ◦ Tθ)(z) − V ∗y∥ 2 Trained Surrogate Latent Representation Tθ : R dr → R dr Latent Space Transport Map Whitened latent prior on R dr Sample z ∼ π = N (0, IdRdr ) Latent representation transport map training min θ Ez∼π h 1 … view at source ↗
Figure 2
Figure 2. Overview of the latent representation lazy map construction. [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗
Figure 3
Figure 3. Overview of LazyDINO amortization procedure. Comparison of LazyDINO to other SBAI methods. Given joint samples of the latent prior and simulated observational data, the SBVI methods optimize for a conditional transport map that matches the pullback distributions, y (j) 7→ Tθ(y (j) , ·) ♯π, to posteriors at the simulated data samples using an fKL objective. The sample generation and the transport map construction are… view at source ↗
Figures from the paper (22 more)
Figure 4
Figure 4. Figure 4: Example I. Setup for inferring the diffusivity field in a nonlinear reaction–diffusion PDE detailed in Section 5.1. For each BIP (#1–4), we show the data-generating synthetic parameter m drawn from the prior, the synthetic data y placed on top of the PDE solution at m,…
Figure 5
Figure 5. Figure 5: Example II. Setup for inferring a heterogeneous hyperelastic material property detailed in Section 5.2. For each BIP (#1–4), we visualize the synthetic parameter m drawn from the prior, the corresponding deformed configuration, the synthetic displacement data y, and th…
Figure 6
Figure 6. Figure 6: Example I neural ridge function surrogate testing. Percentage accuracy, 100% × (1 − error), is also overlaid. Overall, these results demonstrate a significant cost reduction for achieving any given generalization accuracy in both the PtO map and the latent Jacobian via…
Figure 7
Figure 7. Figure 7: Example II neural operator ridge function generalization. In a similar pattern as [PITH_FULL_IMAGE:figures/full_fig_p026_7.png]
Figure 8
Figure 8. Figure 8: Example I moment discrepancies. In all cases, a lower error indicates a better posterior approximation. (LazyDINO vs. LazyNO): Apart from the statistical anomaly at 500 samples attributed to the stochasticity in surrogate training seen in [PITH_FULL_IMAGE:figures/full…
Figure 9
Figure 9. Figure 9: Example II moment discrepancies. We observe similar trends as in Example I ( [PITH_FULL_IMAGE:figures/full_fig_p029_9.png]
Figure 10
Figure 10. Figure 10: Example I density-based diagnostics. Higher values of ESS100K% and lower values of all other diagnostics imply better posterior approximation. We observe similar trends to those observed in the moment discrepancy comparisons in [PITH_FULL_IMAGE:figures/full_fig_p030_…
Figure 11
Figure 11. Figure 11: Example II density-based diagnostics. We observe similar trends as in Example I. Notably, LazyDINO achieves nearly 50% ANIS effective sample percentage as a 100,000-sample independent sampler with 16,000 training samples, 50 to 50,000 times the percentage achieved for…
Figure 12
Figure 12. Figure 12: Example I: LazyDINO vs Laplace approximation marginals. At 250 training samples, we see clear deviation in the marginals for both approaches, though LA-baseline has contours that match the posterior marginals more closely. By 2k training samples, LazyDINO seems to mat…
Figure 13
Figure 13. Figure 13: Example I: LazyDINO vs LazyNO marginals. Consistent with [PITH_FULL_IMAGE:figures/full_fig_p035_13.png]
Figure 14
Figure 14. Figure 14: Example I: LazyDINO vs SBAI marginals. The marginals produced via SBAI are much further from the true posterior marginals than the other approaches. Notably, SBAI consistently overestimates the uncertainty in posterior reconstruction and still yields a poor reconstruc…
Figure 15
Figure 15. Figure 15: Example I: LazyDINO vs LazyMap marginals. The LazyMap marginals are quite poor for 1k training samples and only get more concentrated as the number of training samples increases. The equivalent sample-cost LazyDINO posterior marginals match the ground truth substantia…
Figure 16
Figure 16. Figure 16: Example II: LazyDINO vs Laplace approximation marginals. Due to the non-Gaussianity of the posterior, Laplace approximation struggles to capture the overall behavior of the marginals depicted, whereas the LazyDINO marginals are close to the posterior marginals at 2k t…
Figure 17
Figure 17. Figure 17: Example II: LazyDINO vs LazyNO marginals. With 250 training samples, LazyNO produces far underconcentrated samples in these marginals. With 16k training samples, both methods produce similar marginal distributions. The notable takeaway is that LazyNO requires multiple…
Figure 18
Figure 18. Figure 18: Example II: LazyDINO vs SBAI marginals. SBAI produces samples that are highly under-concentrated. SBAI is simultaneously off in capturing the peak locations of the marginals while overestimating the uncertainty in the parameter. 1 x 2 2 −0.25 0.00 0.25 x 3 0 2 x 4 0.2…
Figure 19
Figure 19. Figure 19: Example II: LazyDINO vs. LazyMap marginals. LazyMap fails to capture the contours of the posterior marginals faithfully. LazyMap is off by a constant error in capturing the peak of the marginals. As more data is available, it tends to over-concentrate, leading to na¨ı…
Figure 20
Figure 20. Figure 20: Example I: progression of mean estimators. LazyDINO already visually captures the mean well with 250 training samples and leads to significantly smaller point-wise errors at 16, 000 samples. The next best performing method is LazyNO, which struggles to resolve the ess…
Figure 21
Figure 21. Figure 21: Example I: progression of MAP estimators. In a similar story to [PITH_FULL_IMAGE:figures/full_fig_p040_21.png]
Figure 22
Figure 22. Figure 22: Example I: progression of point-wise marginal variance estimators. Again, the LazyDINO estimate of point￾wise marginal variance is superior to the other methods as in the previous studies of mean and MAP estimation. Notably, SBAI overestimates the point-wise marginal …
Figure 23
Figure 23. Figure 23: Example II: progression of mean estimators with N. As with [PITH_FULL_IMAGE:figures/full_fig_p042_23.png]
Figure 24
Figure 24. Figure 24: Example II: progression of MAP estimators. As with [PITH_FULL_IMAGE:figures/full_fig_p043_24.png]
Figure 25
Figure 25. Figure 25: Example II: progression of point-wise marginal variance estimators. We see, yet again, the same trend of the previous studies: LazyDINO yields a superior approximation of the point-wise marginal variance and, notably, can deliver a faithful approximation for as little…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

125 extracted references · 61 canonical work pages

  1. [1]

    A. M. Stuart, Inverse problems: A Bayesian perspective, Acta Numerica 19 (2010) 451–559

  2. [2]

    Bui-Thanh, O

    T. Bui-Thanh, O. Ghattas, J. Martin, G. Stadler, A computational framework for infinite-dimensional bayesian inverse problems. part i: The linearized case, with application to global seismic inversion, 2013

  3. [3]

    Petra, J

    N. Petra, J. Martin, G. Stadler, O. Ghattas, A computational framework for infinite-dimensional bayesian inverse problems, part ii: Stochastic newton mcmc with application to ice sheet flow inverse problems, SIAM Journal on Scientific Computing 36 (2014) A1525–A1555

  4. [4]

    Ghattas, K

    O. Ghattas, K. Willcox, Learning physics-based models from data: perspectives from inverse problems and model reduction, Acta Numerica 30 (2021) 445–554

  5. [5]

    M. G. Kapteyn, J. V. Pretorius, K. E. Willcox, A probabilistic graphical model foundation for enabling predictive digital twins at scale, Nature Computational Science 1 (2021) 337–347

  6. [6]

    X. Huan, J. Jagalur, Y. Marzouk, Optimal experimental design: Formulations and computations, Acta Numerica 33 (2024) 715–840

  7. [7]

    Baptista, Y

    R. Baptista, Y. Marzouk, O. Zahm, On the representation and learning of monotone triangular transport maps, Foun- dations of Computational Mathematics (2023)

  8. [8]

    Q. Liu, D. Wang, Stein variational gradient descent: A general purpose bayesian inference algorithm, 2019

Show all 125 references
  1. [9]

    P. Chen, K. Wu, J. Chen, T. O’Leary-Roseberry, O. Ghattas, Projected stein variational newton: A fast and scalable bayesian inference method in high dimensions, Advances in Neural Information Processing Systems 32 (2019)

  2. [10]

    Detommaso, T

    G. Detommaso, T. Cui, A. Spantini, Y. Marzouk, R. Scheichl, A Stein variational Newton method, 2018

  3. [11]

    D. J. Rezende, S. Mohamed, Variational inference with normalizing flows, arXiv:1505.05770 (2015)

  4. [12]

    Papamakarios, E

    G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mohamed, B. Lakshminarayanan, Normalizing flows for probabilistic modeling and inference, arXiv preprint arXiv:1912.02762 (2019)

  5. [13]

    M. C. Brennan, D. Bigoni, O. Zahm, A. Spantini, Y. Marzouk, Greedy inference with structure-exploiting lazy maps, 2020

  6. [14]

    T. Cui, J. Martin, Y. M. Marzouk, A. Solonen, A. Spantini, Likelihood-informed dimension reduction for nonlinear inverse problems, Inverse Problems 30 (2014) 114015

  7. [15]

    P. G. Constantine, Active Subspaces, Society for Industrial and Applied Mathematics, Philadelphia, PA, 2015

  8. [16]

    Bui-Thanh, O

    T. Bui-Thanh, O. Ghattas, Analysis of the Hessian for inverse scattering problems. Part I: Inverse shape scattering of acoustic waves, Inverse Problems 28 (2012) 055001

  9. [17]

    Bui-Thanh, O

    T. Bui-Thanh, O. Ghattas, Analysis of the Hessian for inverse scattering problems. Part II: Inverse medium scattering of acoustic waves, Inverse Problems 28 (2012) 055002

  10. [18]

    Bui-Thanh, O

    T. Bui-Thanh, O. Ghattas, Analysis of the Hessian for inverse scattering problems. Part III: Inverse medium scattering of electromagnetic waves, Inverse Problems and Imaging 7 (2013) 1139–1155

  11. [19]

    P. Chen, O. Ghattas, Hessian-based sampling for high-dimensional model reduction, International Journal for Uncertainty Quantification 9 (2019)

  12. [20]

    P. Chen, U. Villa, O. Ghattas, Hessian-based adaptive sparse quadrature for infinite-dimensional bayesian inverse problems, Computer Methods in Applied Mechanics and Engineering 327 (2017) 147–172

  13. [21]

    P. H. Flath, L. C. Wilcox, V. Ak¸ celik, J. Hill, B. van Bloemen Waanders, O. Ghattas, Fast algorithms for Bayesian uncertainty quantification in large-scale linear inverse problems based on low-rank partial Hessian approximations, SIAM Journal on Scientific Computing 33 (2011...

  14. [22]

    Isaac, N

    T. Isaac, N. Petra, G. Stadler, O. Ghattas, Scalable and efficient algorithms for the propagation of uncertainty from data through inference to prediction for large-scale problems, with application to flow of the Antarctic ice sheet, Journal of Computational Physics 296 (2015) 348–368

  15. [23]

    Spantini, A

    A. Spantini, A. Solonen, T. Cui, J. Martin, L. Tenorio, Y. Marzouk, Optimal low-rank approximations of bayesian linear inverse problems, SIAM Journal on Scientific Computing 37 (2015) A2451–A2487

  16. [24]

    O’Leary-Roseberry, U

    T. O’Leary-Roseberry, U. Villa, P. Chen, O. Ghattas, Derivative-informed projected neural networks for high-dimensional parametric maps governed by PDEs, Computer Methods in Applied Mechanics and Engineering 388 (2022) 114199

  17. [25]

    Hesthaven, S

    J. Hesthaven, S. Ubbiali, Non-intrusive reduced order modeling of nonlinear problems using neural networks, Journal of Computational Physics 363 (2018) 55–78

  18. [26]

    Kovachki, Z

    N. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. Stuart, A. Anandkumar, Neural operator: Learning maps between function spaces with applications to PDEs, Journal of Machine Learning Research 24 (2023) 1–97. 46

  19. [27]

    N. B. Kovachki, S. Lanthaler, A. M. Stuart, Operator learning: Algorithms and analysis, arxiv.2402.15715 (2024)

  20. [28]

    D. Luo, T. O’Leary-Roseberry, P. Chen, O. Ghattas, Efficient PDE-constrained optimization under high-dimensional uncertainty using derivative-informed neural operators, arXiv preprint arXiv:2305.20053 (2023)

  21. [29]

    L. Cao, T. O’Leary-Roseberry, O. Ghattas, Derivative-informed neural operator acceleration of geometric MCMC for infinite-dimensional Bayesian inverse problems, arXiv preprint arXiv:2403.08220 (2024)

  22. [30]

    O’Leary-Roseberry, P

    T. O’Leary-Roseberry, P. Chen, U. Villa, O. Ghattas, Derivative-informed neural operator: an efficient framework for high-dimensional parametric derivative learning, Journal of Computational Physics 496 (2024) 112555

  23. [31]

    T. Cui, O. Zahm, Data-free likelihood-informed dimension reduction of Bayesian inverse problems, Inverse Problems 37 (2021) 045009

  24. [32]

    O. Zahm, P. G. Constantine, C. Prieur, Y. M. Marzouk, Gradient-based dimension reduction of multivariate vector-valued functions, SIAM Journal on Scientific Computing 42 (2020) A534–A558

  25. [33]

    P. S. Laplace, Memoir on the probability of the causes of events (1774). m´ emoires de math´ ematique et de physique, tome sixi` eme. (english translation by s. m. stigler), Statistical science 1 (1986) 364–378

  26. [34]

    Bui-Thanh, C

    T. Bui-Thanh, C. Burstedde, O. Ghattas, J. Martin, G. Stadler, L. C. Wilcox, Extreme-scale uq for bayesian inverse prob- lems governed by pdes, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’12, IEEE Compute...

  27. [35]

    Y. M. Marzouk, H. N. Najm, Dimensionality reduction and polynomial chaos acceleration of bayesian inference in inverse problems, Journal of Computational Physics 228 (2009) 1862–1902

  28. [36]

    O. Zahm, T. Cui, K. Law, A. Spantini, Y. Marzouk, Certified dimension reduction in nonlinear Bayesian inverse problems, Mathematics of Computation 91 (2022) 1789–1835

  29. [37]

    Bigoni, Y

    D. Bigoni, Y. Marzouk, C. Prieur, O. Zahm, Nonlinear dimension reduction for surrogate modeling using gradient information, Information and Inference: A Journal of the IMA 11 (2022) 1597–1639

  30. [38]

    Y. M. Marzouk, H. N. Najm, L. A. Rahn, Stochastic spectral methods for efficient Bayesian solution of inverse problems, Journal of Computational Physics 224 (2007) 560–586

  31. [39]

    Marzouk, D

    Y. Marzouk, D. Xiu, A stochastic collocation approach to bayesian inference in inverse problems, COMMUNICATIONS IN COMPUTATIONAL PHYSICS 6 (2009) 826–847

  32. [40]

    Farcas, J

    I.-G. Farcas, J. Latz, E. Ullmann, T. Neckel, H.-J. Bungartz, Multilevel adaptive sparse Leja approximations for Bayesian inverse problems, SIAM Journal on Scientific Computing 42 (2020) A424–A451

  33. [41]

    Galbally, K

    D. Galbally, K. Fidkowski, K. Willcox, O. Ghattas, Non-linear model reduction for uncertainty quantification in large- scale inverse problems, International Journal for Numerical Methods in Engineering 81 (2010) 1581–1608

  34. [42]

    Lieberman, K

    C. Lieberman, K. Willcox, O. Ghattas, Parameter and state model reduction for large-scale statistical inverse problems, SIAM Journal on Scientific Computing 32 (2010) 2523–2542

  35. [43]

    T. Cui, Y. M. Marzouk, K. E. Willcox, Data-driven model reduction for the Bayesian solution of inverse problems, International Journal for Numerical Methods in Engineering 102 (2015) 966–990

  36. [44]

    Peherstorfer, K

    B. Peherstorfer, K. Willcox, M. Gunzburger, Survey of multifidelity methods in uncertainty propagation, inference, and optimization, SIAM Review 60 (2018) 550–591

  37. [45]

    M. B. Lykkegaard, T. J. Dodwell, C. Fox, G. Mingas, R. Scheichl, Multilevel delayed acceptance MCMC, SIAM/ASA Journal on Uncertainty Quantification 11 (2023) 1–30

  38. [46]

    L. Cao, T. O’Leary-Roseberry, P. K. Jha, J. T. Oden, O. Ghattas, Residual-based error correction for neural operator accelerated infinite-dimensional Bayesian inverse problems, Journal of Computational Physics 486 (2023) 112104

  39. [47]

    Bhattacharya, B

    K. Bhattacharya, B. Hosseini, N. B. Kovachki, A. M. Stuart, Model reduction and neural network for parametric PDEs, The SMAI Journal of computational mathematics 7 (2021) 121–157

  40. [48]

    Fresca, A

    S. Fresca, A. Manzoni, POD-DL-ROM: Enhancing deep learning-based reduced order models for nonlinear parametrized PDEs by proper orthogonal decomposition, Computer Methods in Applied Mechanics and Engineering 388 (2022) 114181

  41. [49]

    O’Leary-Roseberry, X

    T. O’Leary-Roseberry, X. Du, A. Chaudhuri, J. R. Martins, K. Willcox, O. Ghattas, Learning high-dimensional para- metric maps via reduced basis adaptive residual networks, Computer Methods in Applied Mechanics and Engineering 402 (2022) 115730

  42. [50]

    L. Lu, P. Jin, G. Pang, Z. Zhang, G. E. Karniadakis, Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Nature Machine Intelligence 3 (2021) 218–229

  43. [51]

    J. H. Seidman, G. Kissas, G. J. Pappas, P. Perdikaris, Variational autoencoding neural operators, arXiv preprint, arXiv.2302.10351 (2023)

  44. [52]

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anandkumar, Fourier neural operator for parametric partial differential equations, arXiv preprint, arXiv.2010.08895 (2021)

  45. [53]

    Q. Cao, S. Goswami, G. E. Karniadakis, LNO: Laplace neural operator for solving differential equations, arXiv preprint, arXiv.2303.10528 (2023)

  46. [54]

    Lanthaler, Z

    S. Lanthaler, Z. Li, A. M. Stuart, The nonlocal neural operator: Universal approximation, arXiv preprint, arXiv.2304.13221 (2023)

  47. [55]

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anandkumar, Neural operator: Graph kernel network for partial differential equations, arXiv preprint, arXiv.2003.03485 (2020)

  48. [56]

    Z. Li, H. Zheng, N. Kovachki, D. Jin, H. Chen, B. Liu, K. Azizzadenesheli, A. Anandkumar, Physics-informed neural operator for learning partial differential equations, ACM / IMS Journal of Data Science (2024)

  49. [57]

    S. Wang, H. Wang, P. Perdikaris, Learning the solution operator of parametric partial differential equations with physics- informed DeepONets, Science Advances 7 (2021) eabi8605

  50. [58]

    J. Go, P. Chen, Accelerating Bayesian Optimal Experimental Design with Derivative-Informed Neural Operators, arXiv preprint arXiv:2312.14810 (2023). 47

  51. [59]

    J. Go, P. Chen, Sequential infinite-dimensional Bayesian optimal experimental design with derivative-informed latent attention neural operator, arXiv preprint arXiv:2409.09141 (2024)

  52. [60]

    Y. Qiu, N. Bridges, P. Chen, Derivative-enhanced deep operator network, arXiv.2402.19242 (2024)

  53. [61]

    E. G. Tabak, C. V. Turner, A family of nonparametric density estimation algorithms, Communications on Pure and Applied Mathematics 66 (2013) 145–164

  54. [62]

    Kobyzev, S

    I. Kobyzev, S. Prince, M. Brubaker, Normalizing flows: An introduction and review of current methods, IEEE Transac- tions on Pattern Analysis and Machine Intelligence (2020)

  55. [63]

    L. Dinh, J. Sohl-Dickstein, S. Bengio, Density estimation using real NVP, arXiv:1605.08803 (2016)

  56. [64]

    D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, M. Welling, Improved variational inference with inverse autoregressive flow, in: D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, R. Garnett (Eds.), Advances in Neural Information Processing Systems 29, Curran As...

  57. [65]

    Papamakarios, T

    G. Papamakarios, T. Pavlakou, I. Murray, Masked autoregressive flow for density estimation, in: Advances in Neural Information Processing Systems, pp. 2338–2347

  58. [66]

    Huang, D

    C.-W. Huang, D. Krueger, A. Lacoste, A. Courville, Neural autoregressive flows, in: International Conference on Machine Learning, PMLR, pp. 2078–2087

  59. [67]

    De Cao, I

    N. De Cao, I. Titov, W. Aziz, Block neural autoregressive flow, arXiv preprint arXiv:1904.04676 (2019)

  60. [68]

    Jaini, K

    P. Jaini, K. A. Selby, Y. Yu, Sum-of-squares polynomial flow, in: International Conference on Machine Learning, pp. 3009–3018

  61. [69]

    Daniels, M

    H. Daniels, M. Velikova, Monotone and partially monotone neural networks, IEEE Transactions on Neural Networks 21 (2010) 906–917

  62. [70]

    Wehenkel, G

    A. Wehenkel, G. Louppe, Unconstrained monotonic neural networks, Advances in neural information processing systems 32 (2019)

  63. [71]

    Knothe, Contributions to the theory of convex bodies., Michigan Mathematical Journal 4 (1957) 39–52

    H. Knothe, Contributions to the theory of convex bodies., Michigan Mathematical Journal 4 (1957) 39–52

  64. [72]

    Rosenblatt, Remarks on a Multivariate Transformation, The Annals of Mathematical Statistics 23 (1952) 470 – 472

    M. Rosenblatt, Remarks on a Multivariate Transformation, The Annals of Mathematical Statistics 23 (1952) 470 – 472

  65. [73]

    T. A. El Moselhy, Y. M. Marzouk, Bayesian inference with optimal maps, Journal of Computational Physics 231 (2012) 7815–7850

  66. [74]

    Marzouk, T

    Y. Marzouk, T. Moselhy, M. Parno, A. Spantini, Sampling via measure transport: An introduction, in: Handbook of Uncertainty Quantification, R. Ghanem, D. Higdon, and H. Owhadi, editors, Springer, 2016

  67. [75]

    Spantini, D

    A. Spantini, D. Bigoni, Y. Marzouk, Inference via low-dimensional couplings, The Journal of Machine Learning Research 19 (2018) 2639–2709

  68. [76]

    Baptista, O

    R. Baptista, O. Zahm, Y. Marzouk, An adaptive transport framework for joint and conditional density estimation, arXiv:2009.10303 (2020)

  69. [77]

    J. Zech, Y. Marzouk, Sparse approximation of triangular transports, Part II: The infinite-dimensional case, Constr. Approx. 55 (2022) 987–1036

  70. [78]

    Westermann, J

    J. Westermann, J. Zech, Measure transport via polynomial density surrogates, arXiv preprint (2023)

  71. [79]

    J. Zech, Y. Marzouk, Sparse approximation of triangular transports, Part I: The finite-dimensional case, Constr. Approx. 55 (2022) 919–986

  72. [80]

    Zeghal, F

    J. Zeghal, F. Lanusse, A. Boucaud, B. Remy, E. Aubourg, Neural posterior estimation with differentiable simulators, 2022

  73. [81]

    Brehmer, G

    J. Brehmer, G. Louppe, J. Pavez, K. Cranmer, Mining gold from implicit models to improve likelihood-free inference, Proceedings of the National Academy of Sciences 117 (2020) 5242–5249

  74. [82]

    Durkan, G

    C. Durkan, G. Papamakarios, I. Murray, Sequential neural methods for likelihood-free inference, arXiv preprint arXiv:1811.08723 (2018)

  75. [83]

    Papamakarios, D

    G. Papamakarios, D. Sterratt, I. Murray, Sequential neural likelihood: Fast likelihood-free inference with autoregressive flows, in: The 22nd International Conference on Artificial Intelligence and Statistics, PMLR, pp. 837–848

  76. [84]

    Greenberg, M

    D. Greenberg, M. Nonnenmacher, J. Macke, Automatic posterior transformation for likelihood-free inference, in: Inter- national Conference on Machine Learning, PMLR, pp. 2404–2414

  77. [85]

    Papamakarios, I

    G. Papamakarios, I. Murray, Fast ϵ-free inference of simulation models with bayesian conditional density estimation, 2018

  78. [86]

    Baptista, L

    R. Baptista, L. Cao, J. Chen, O. Ghattas, F. Li, Y. M. Marzouk, J. T. Oden, Bayesian model calibration for block copolymer self-assembly: Likelihood-free inference and expected information gain computation via measure transport, Journal of Computational Physics 503 (2024) 112844

  79. [87]

    Ganguly, S

    A. Ganguly, S. Jain, U. Watchareeruetai, Amortized variational inference: A systematic review, Journal of Artificial Intelligence Research 78 (2023) 167–215

  80. [88]

    Soize, R

    C. Soize, R. Ghanem, Physical systems with random uncertainties: Chaos representations with arbitrary probability measure, SIAM Journal on Scientific Computing 26 (2004) 395–410

  81. [89]

    Coordinate transformation and polynomial chaos for the bayesian inference of a gaussian process with parametrized prior covariance function, Computer Methods in Applied Mechanics and Engineering 298 (2016) 205–228

  82. [90]

    D. M. Blei, A. Kucukelbir, J. D. McAuliffe, Variational inference: A review for statisticians, Journal of the American statistical Association 112 (2017) 859–877

  83. [91]

    D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, M. Welling, Improving variational inference with inverse autoregressive flow, 2017

  84. [92]

    Bottou, F

    L. Bottou, F. E. Curtis, J. Nocedal, Optimization methods for large-scale machine learning, SIAM review 60 (2018) 223–311

  85. [93]

    Isaac, N

    T. Isaac, N. Petra, G. Stadler, O. Ghattas, Scalable and efficient algorithms for the propagation of uncertainty from 48 data through inference to prediction for large-scale problems, with application to flow of the antarctic ice sheet, Journal of Computational Physics 296 (20...

  86. [94]

    Villa, N

    U. Villa, N. Petra, O. Ghattas, hIPPYlib: An extensible software framework for large-scale inverse problems governed by PDEs: Part I: Deterministic inversion and linearized Bayesian inference, ACM Transactions on Mathematical Software 47 (2021)

  87. [95]

    Villa, T

    U. Villa, T. O’Leary-Roseberry, A note on the relationship between pde-based precision operators and mat \’ern covari- ances, arXiv preprint arXiv:2407.00471 (2024)

  88. [96]

    Kirchhoff, D

    J. Kirchhoff, D. Luo, T. O’Leary-Roseberry, O. Ghattas, Inference of Heterogeneous Material Properties via Infinite- Dimensional Integrated DIC, arXiv preprint arXiv:2408.10217 (2024)

  89. [97]

    D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, 2017

  90. [98]

    Beskos, G

    A. Beskos, G. Roberts, A. Stuart, J. Voss, MCMC methods for diffusion bridges, Stochastics and Dynamics 8 (2008) 319–350

  91. [99]

    Agapiou, O

    S. Agapiou, O. Papaspiliopoulos, D. Sanz-Alonso, A. M. Stuart, Importance sampling: Intrinsic dimension and compu- tational cost, 2017

  92. [100]

    B. Li, T. Bengtsson, P. Bickel, Curse-of-dimensionality revisited: Collapse of importance sampling in very large scale systems (2005)

  93. [101]

    F. M. Polo, R. Vicente, Effective sample size, dimensionality, and generalization in covariate shift adaptation, Neural Computing and Applications 35 (2020) 18187–18199

  94. [102]

    Sanz-Alonso, Z

    D. Sanz-Alonso, Z. Wang, Bayesian update with importance sampling: Required sample size, Entropy 23 (2020) 22

  95. [103]

    Beskos, M

    A. Beskos, M. Girolami, S. Lan, P. E. Farrell, A. M. Stuart, Geometric MCMC for infinite-dimensional inverse problems, Journal of Computational Physics 335 (2017) 327–351

  96. [104]

    Nualart, The Malliavin calculus and related topics, volume 1995, Springer, 2006

    D. Nualart, The Malliavin calculus and related topics, volume 1995, Springer, 2006

  97. [105]

    V. I. Bogachev, Gaussian measures, 62, American Mathematical Soc., 1998

  98. [106]

    Halko, P.-G

    N. Halko, P.-G. Martinsson, J. A. Tropp, Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions, 2010

  99. [107]

    Xiang, J

    H. Xiang, J. Zou, Randomized algorithms for large-scale inverse problems with general regularizations, 2014

  100. [108]

    A. K. Saibaba, J. Lee, P. K. Kitanidis, Randomized algorithms for generalized hermitian eigenvalue problems with application to computing karhunen-lo` eve expansion, 2015

  101. [109]

    G. H. Golub, Q. Ye, An inverse free preconditioned krylov subspace method for symmetric generalized eigenvalue problems, SIAM Journal on Scientific Computing 24 (2002) 312–334

  102. [110]

    Saad, Krylov subspace methods for solving large unsymmetric linear systems, Mathematics of Computation 37 (1981) 105–126

    Y. Saad, Krylov subspace methods for solving large unsymmetric linear systems, Mathematics of Computation 37 (1981) 105–126

  103. [111]

    D. C. Sorensen, Truncated qz methods for large scale generalized eigenvalue problems, Electron. Trans. Numer. Anal. 7 (1998) 141–162

  104. [112]

    van den Eshof, G

    J. van den Eshof, G. L. G. Sleijpen, Inexact krylov subspace methods for linear systems, SIAM Journal on Matrix Analysis and Applications 26 (2004) 125–153

  105. [113]

    Chowdhury, The truncated lanczos algorithm for partial solution of the symmetric eigenproblem, Computers & Structures 6 (1976) 439–446

    P. Chowdhury, The truncated lanczos algorithm for partial solution of the symmetric eigenproblem, Computers & Structures 6 (1976) 439–446

  106. [114]

    Z. Yao, A. Gholami, S. Shen, M. Mustafa, K. Keutzer, M. Mahoney, Adahessian: An adaptive second order optimizer for machine learning, in: proceedings of the AAAI conference on artificial intelligence, volume 35, pp. 10665–10673

  107. [115]

    O’Leary-Roseberry, R

    T. O’Leary-Roseberry, R. Bollapragada, Fast Unconstrained Optimization via Hessian Averaging and Adaptive Gradient Sampling Methods, arXiv preprint arXiv:2408.07268 (2024)

  108. [116]

    Bollapragada, R

    R. Bollapragada, R. Byrd, J. Nocedal, Adaptive sampling strategies for stochastic optimization, SIAM Journal on Optimization 28 (2018) 3312–3343

  109. [117]

    Newton, R

    D. Newton, R. Bollapragada, R. Pasupathy, N. K. Yip, A retrospective approximation approach for smooth stochastic optimization, 2024

  110. [118]

    Kretschmann, Are Minimizers of the Onsager–Machlup Functional Strong Posterior Modes?, SIAM/ASA Journal on Uncertainty Quantification 11 (2023) 1105–1138

    R. Kretschmann, Are Minimizers of the Onsager–Machlup Functional Strong Posterior Modes?, SIAM/ASA Journal on Uncertainty Quantification 11 (2023) 1105–1138

  111. [119]

    R. S. Dembo, S. C. Eisenstat, T. Steihaug, Inexact newton methods, SIAM Journal on Numerical Analysis 19 (1982) 400–408

  112. [120]

    S. C. Eisenstat, H. F. Walker, Choosing the forcing terms in an inexact newton method, SIAM Journal on Scientific Computing 17 (1996) 16–32. Appendix A. Glossary of terminology amortized cost: Discounted cost, by spreading it out over the solution of additional problems. The m...

  113. [121]

    , N ▷ Sample prior

    m(j) ∼ µ, j = 1, . . . , N ▷ Sample prior

  114. [122]

    , N ▷ Evaluate PtO map/Jacobian

    G(m(j)), DH G(m(j)), j = 1, . . . , N ▷ Evaluate PtO map/Jacobian

  115. [123]

    Create encoder/decoder: {ψk ∈ M }dr k=1 ← eigenvalue problem( DG(m(j)) NL i=1, Γ−1 n , C−1, ϵL or dr) ▷ (32), (E.1) Drz := Pdr k=1 zkψk, Er := D⊤ r C−1 ▷ (15), (17)

  116. [124]

    , N end Latent space: solving the generalized eigenvalue problem

    Embed dataset: ▷ (38), (42) z(j) ← Erm(j), g(j) ← V ⊤Γ−1 n G(m(j)), J (j) r ← V ⊤Γ−1 n DG(m(j))Dr, j = 1, . . . , N end Latent space: solving the generalized eigenvalue problem. We describe here the computation of the eigenvalue problem in Algorithm 1 to find the reduced basis...

  117. [125]

    Appendix E.2

    Train gw by minimizing an empirical risk: w∗ = argmin w∈RdW 1 N PN j=1 g(j) − gw z(j) 2 + J (j) r − ∇zgw z(j) 2 F| {z } include for H 1µ (RB-DINO) objective end Equipped with a sufficiently accurate neural network approximation to the optimal latent PtO map gopt, including acc...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.