Pith. sign in

REVIEW 4 major objections 6 minor 56 references

Diffusion-based annealed Boltzmann generators gain from second-order or deterministic transport-map corrections, but practical failures trace to learned log-density error, not score error.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Even a perfect diffusion model yields poor annealed Boltzmann generators when coupled through first-order stochastic denoising kernels, while deterministic transport maps and second-order kernels improve; with learned densities, log-density error, not score error, is the bottleneck.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection A useful idealized-regime decomposition and a promising deterministic transport variant, but the unbiased log-det claim is overstated and the 'systematic failure' headline runs ahead of the evidence. the 4 major comments →

arxiv 2601.21026 v2 pith:YWZZUFRD submitted 2026-01-28 stat.ML cs.LG

Diffusion-based Annealed Boltzmann Generators : benefits, pitfalls and hopes

classification stat.ML cs.LG MSC 65C05
keywords Boltzmann generatorsdiffusion modelsannealed Monte Carlosequential Monte Carloreplica exchangelog-density estimationmode blindnesstransport maps
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether diffusion models can serve as the backbone of Boltzmann Generators by providing the intermediate-density path for annealed Monte Carlo. In an idealized setting with a perfectly known diffusion model, it shows that the diffusion density path beats classic tempering paths, but that standard first-order stochastic denoising kernels give no improvement over a naive baseline. Second-order kernels that use covariance information substantially improve performance, and a new deterministic transport-map construction nearly matches them without needing covariance estimates. In a realistic setting with a learned diffusion model, all annealed variants fail, and the paper argues the bottleneck is inaccurate learned log-densities, not scores.

Core claim

The systematic empirical study isolates inference effects from learning effects by comparing a perfectly learned diffusion model with one trained from data. With exact scores and log-densities, first-order stochastic denoising kernels—which only match the conditional mean—perform no better than a correlation-free baseline, whereas second-order Gaussian kernels that incorporate conditional covariance yield large gains. The paper then introduces deterministic transitions derived from the probability-flow ODE, built with an implicit midpoint integrator whose forward and backward maps are mutual inverses; estimating the Jacobian log-determinants via a power series and the Hutchinson trace trick

What carries the argument

The central object is the diffusion-induced density path together with the transition kernels/maps between adjacent noise levels: first-order stochastic denoising kernels (score-only), second-order stochastic kernels using Hessian-based covariance, and the newly proposed deterministic implicit-midpoint integrators of the probability-flow ODE, whose mutual invertibility and power-series Jacobian log-determinants (with Hutchinson trace estimation) make them usable inside AIS, SMC, and replica-exchange annealed samplers.

Load-bearing premise

The conclusion that log-density error, not score error, is the bottleneck assumes that the gap between the well-performing learned reverse dynamics and the failing annealed samplers is entirely due to log-density inaccuracy, and that the trained energy-based architectures and losses tested are representative of diffusion log-density estimation; if the hardcoded scores are imperfect at low temperature or the failures come from capacity or training instability rather than mode

What would settle it

Train a diffusion log-density model with an objective that provably recovers mode proportions (e.g., component-wise reweighting or a mode-aware regularizer) on the 16-mode Gaussian mixture, then rerun the AIS/SMC/RE comparisons; if performance jumps to idealized levels, the mode-blindness bottleneck is confirmed, and if not, it is falsified. For the idealized hierarchy, compute effective sample sizes for first-order versus second-order kernels on a two-mode Gaussian mixture with known exact conditional covariance; if first-order matches second-order, the claim that first-order kernels fail wou

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Diffusion density paths should replace tempering paths in annealed samplers for multimodal targets, since they avoid abrupt mode switching and preserve relative mode weights.
  • First-order stochastic denoising kernels are not worth the extra computation in diffusion-based annealed Boltzmann generators; second-order or deterministic transport-map corrections are required for meaningful gains.
  • The deterministic transport-map framework provides a practical alternative to second-order methods, achieving comparable accuracy with only score information and a modest computational overhead.
  • In realistic settings, score accuracy is insufficient: annealed samplers fail because learned log-densities misrepresent mode proportions, even when the learned reverse dynamics are accurate.
  • Training objectives that suffer from mode blindness will systematically disrupt SMC resampling and replica-exchange communication on multi-modal targets, dominating any benefit from improved transitions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The Hutchinson-based log-determinant estimation could be adapted to other flow-based or transport-based samplers, offering an unbiased acceptance correction without explicit Hessians.
  • The diagnosis points research toward log-density estimators that enforce correct mode weights; if such training schemes are developed, iterative diffusion-based annealed samplers could become viable.
  • The mode-blindness mechanism likely affects any diffusion-based SMC or inference-time alignment algorithm that resamples using learned log-densities, even when the score is well learned.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper studies diffusion-model-based annealed Monte Carlo Boltzmann Generators (DM-aMC-BGs) on controlled Gaussian-mixture targets, separating an idealized regime (exact scores and log-densities) from a realistic regime (learned energy-based parameterizations). It reports three main findings: (i) diffusion density paths generally outperform tempering paths in aMC; (ii) in the idealized regime, first-order stochastic denoising kernels give little or no improvement over a correlation-free baseline, while second-order stochastic kernels and a newly proposed deterministic transport-map integrator give substantial gains; and (iii) in the learned regime all DM-aMC-BG variants degrade, with the paper attributing the failure primarily to inaccurate, mode-blind DM log-density estimates rather than to learned scores. The paper includes a large appendix with proofs, additional experiments, and ablations, and the code is publicly available.

Significance. If its central claims hold, the paper makes a useful contribution: it provides a controlled benchmark that cleanly separates inference error from learning error, it identifies a concrete limitation of first-order stochastic denoising kernels inside aMC, it proposes a deterministic alternative that may be of independent interest, and it formulates a falsifiable hypothesis about mode blindness in DM log-density estimation. The strengths are the extensive idealized-regime experiments on Gaussian mixtures, the reproducible code release, and the unusually transparent discussion of limitations. However, the statistical-guarantee claim for the deterministic transport-map estimator is currently not correct as stated, and the empirical claims are stronger than the evidence presented in the figures. These issues bear directly on the headline conclusions and need to be addressed.

major comments (4)
  1. [§4.2, Prop. 3] The proposed Jacobian log-determinant estimator is not unbiased. The estimator truncates the power series at order I and replaces the implicit maps with finite-M fixed-point iterates. The Hutchinson estimator gives an unbiased estimate of each truncated trace term Tr([A^(M)]^i), not of log|det J|; no Russian-roulette debiasing is used. Therefore the log-det estimate has bias from truncation and from fixed-point error, and the importance weights in (21) and the acceptance probabilities in (22) are biased. The invocation of Andrieu & Roberts (2009) is not justified: pseudo-marginal MH requires an unbiased estimate of the weight/acceptance ratio, not of its logarithm. Since the central positive claim is that the deterministic transport map outperforms stochastic first-order variants, the reported advantage could be an artifact of this bias; the sensitivity to M and I shown in Figures 55-57
  2. [Fig. 3 / §3.3 / Abstract] The empirical support for the strong wording 'fail systematically' is insufficient. Figure 3 and the related figures show only averages over 8 runs with no error bars, standard deviations, or confidence intervals. The text in §3.3 is more cautious — 'do not yield noticeable improvements' — and the figure caption itself states that the first-order stochastic kernels 'does not always lead to better performance' than the baseline. Without uncertainty quantification, one cannot distinguish a systematic failure from statistical noise. Please report error bars or confidence intervals and align the abstract, §3.3, and the introduction with the actual strength of the empirical statement.
  3. [§6.2] The claim that the realistic-regime failure is 'not imputable to the quality of the learned scores, but rather to the learned log-densities' is a comparative inference, not a directly measured quantity. Figure 6 is a qualitative 1D visualization of learned density paths; no quantitative job-level error for learned log-densities versus learned scores is reported. Alternative explanations — for example network capacity, training instability (especially for the pinned architecture), or imperfect scores in low-temperature/tail regions — are not excluded. Since the 'bottleneck is inaccurate DM log-density estimation' message is one of the paper's main takeaways, please add quantitative diagnostics (e.g., mode-weight errors for the learned densities, score-error norms at intermediate levels), or perform ablations that hold one component fixed while varying the other, to isolate the claimed mec
  4. [§4.1-4.2] The finite fixed-point approximation also undermines the exactness of the deterministic aMC formulation itself. The derivation of the AIS weight (21) and the RE acceptance probability (22) relies on the mutual invertibility property (20). With M fixed-point iterations, the maps are not guaranteed to be mutual inverses, and the paper explicitly leaves the rejection-based safeguard to future work. Thus, even setting aside the log-det estimator bias, the implemented procedure is not the exact deterministic aMC described in §4.1. Please state this clearly and quantify the effect of M on the validity of the transport-map construction, or incorporate the rejection step so that the implemented algorithm matches the claimed statistical framework.
minor comments (6)
  1. [Abstract vs. §3.3] The abstract states that 'standard integrations using only first-order stochastic denoising kernels fail systematically,' while §3.3 says they 'do not yield noticeable improvements.' These are different claims; please use consistent wording.
  2. [Proof of Prop. 3, Eq. (34)] In the displayed equations, the second line writes log|det J_{T_{k+1|k}}(x_k)| where it should be log|det J_{T_{k|k+1}}(x_{k+1})|. Please correct the subscript.
  3. [§6.2] Typo: 'Harcoded' should be 'Hardcoded'.
  4. [Section C] The proof section says 'We leave the proof for the reader' for the EI-based VP/VE variants of the main propositions. These variants are used in the experiments, so either provide the proofs or state explicitly that they are direct substitutions and indicate where the required assumptions differ.
  5. [Abstract] The word 'meta-analysis' is unusual for a controlled empirical study with a new method. Consider replacing it with 'empirical study' or 'comparative analysis' to avoid implying a formal meta-analysis of the literature.
  6. [Figures 1 and 3] The figure captions note that darker bars correspond to larger K and that configurations do not share computational budget. This is important context but easy to miss; consider making it more prominent in the main-text discussion when claiming that diffusion paths 'outperform' tempering paths.

Circularity Check

0 steps flagged

No derivation in the paper reduces to its own inputs; the central findings are new controlled experiments. Minor self-citation to the authors' own benchmark targets is present but not load-bearing. A flagged correctness gap in Section 4.2 about the 'unbiased' Jacobian log-determinant estimator is a validity concern, not a circularity.

full rationale

Walking the derivation chain: in the idealized regime (Section 3 and 4), all intermediate log-densities, scores, and Hessians are computed exactly for Gaussian-mixture targets, so no fitted constant is reused as a prediction. The comparisons among first-order stochastic kernels, second-order stochastic kernels, and deterministic transports are empirical measurements on controlled benchmarks, not consequences of an equation that already contains the answer. The claimed advantage of deterministic transport is a new experimental finding, and its sensitivity to truncation/fixed-point hyperparameters is openly ablated (Figures 55–57). In the realistic regime (Section 6), the same network E_theta provides both the learned log-density and the learned score s_theta = -grad E_theta. The paper argues that since the learned reverse SDE/ODE (score-only) matches its ideal analog while the aMC variants (which additionally use E_theta values) fail, the bottleneck is inaccurate log-density estimation. This is a plausible statistical decomposition, but it is not a circular equivalence by construction: the score and the log-density are tied through one architecture, so a trajectory on which the score is accurate does not logically force the log-density level to be accurate elsewhere, and the conclusion depends on the representatives of the trained EBMs. This is a confounding/correctness risk, not a self-definitional reduction. The main self-referential elements are citations to the authors' own previous work for the benchmark targets (Grenioux et al. 2025; Noble et al. 2025) and for a Tweedie-style expansion (Grenioux et al. 2024). These are not used as uniqueness theorems or as smuggled ansatze; the target definitions are simply Gaussian-mixture densities, and the expansion is reproducible from the paper's own Lemmas 4–5 and Corollary 5. Thus the self-citations do not carry the derivational weight of the paper's claims. I flag one explicit validity gap, although it is not circularity: Section 4.2 calls the procedure an 'unbiased estimator, thereby preserving the statistical guarantees of aMC' and invokes Andrieu & Roberts (2009), but Proposition 3 is stated as an 'approximation' with a truncated power series (order I) and a finite fixed-point range (M). The Hutchinson estimator is unbiased for each truncated trace term, not for the untruncated log-determinant of the implicit-midpoint map; no Russian-roulette debiasing is used and the paper explicitly leaves it to future

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The paper introduces a new integration scheme (implicit midpoint deterministic transport maps) and an estimation recipe (truncated power series + Hutchinson traces), but no new physical entity, force, or conserved quantity. The main load-bearing assumptions are the smoothness/small-step condition for the implicit maps and the mode-blindness explanation for learned-density failure; the latter is explicitly conjectural.

free parameters (3)
  • K (number of annealing levels) = optimized per method/target among {16, 32, 64, 128, 256}
    The main comparisons display diffusion results for all K but tempering 'optimized on K'; in Figure 4 each method uses its best K, so K is effectively selected to show best performance, and computational budgets differ across configurations.
  • lambda_reg (regularization weight for LFPE/aLFPE/RNE) = best of {1e-2, 1e-3, 1e-4} for LFPE/aLFPE and {10, 1, 0.1} for RNE
    Section D.2 states that for LFPE/aLFPE/RNE objectives 'we systematically display, for each sampling/target setting, the best sampling results obtained among the considered values of lambda_reg.' This is post hoc selection over a hyperparameter.
  • M, I, NH (fixed-point iterations, truncation order, Hutchinson samples) = M=4, I=3, NH=32
    Selected as a compromise in ablation studies; they are hand-chosen and the truncation at I=3 introduces bias in the claimed-unbiased Jacobian log-determinant estimator.
axioms (4)
  • standard math Reverse-time SDE / PF-ODE formulation of diffusion models (Anderson 1982; Song et al. 2021) and Tweedie's formula for conditional denoising moments.
    Used throughout Section 2.1 and in the derivation of first/second-order kernels (Eqs. 5-10).
  • ad hoc to paper Assumption 1/3: the score function is L_k-Lipschitz and the step size is small enough, δ_k = O(1/L_k), for the implicit midpoint fixed-point iteration to converge.
    This assumption underpins Propositions 1-3 for the new deterministic transport maps, but the main experiments fix M=4 and never verify the Lipschitz/small-step condition on the actual test targets.
  • domain assumption The controlled Gaussian mixture targets (TwoModes, ManyModes), standardized to zero mean and unit covariance, are adequate proxies for the multi-modal, high-barrier sampling problems encountered in molecular simulation.
    The paper explicitly justifies this choice in the introduction and Section 7, arguing methods that fail here are unlikely to scale; this is a domain assumption, not a theorem.
  • ad hoc to paper Mode blindness of score-based objectives (Wenliang & Kanagawa 2021) extends to TSM, tSM, LFPE, aLFPE, and RNE objectives and is the primary cause of the realistic-regime failure.
    Section 6.2 and Section 7 present this as a conjecture ('we conjecture', 'could primarily be due') supported by empirical density-path plots, not by a proof or an intervention that isolates mode blindness.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Diffusion-based Annealed Boltzmann Generators : benefits, pitfalls and hopes." pith.science (2026). https://pith.science/paper/YWZZUFRD

@misc{pith2026260121026,
  author       = {Pith},
  title        = {Pith review of: Diffusion-based Annealed Boltzmann Generators : benefits, pitfalls and hopes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YWZZUFRD}},
  note         = {Machine review of arXiv:2601.21026}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Sampling configurations at thermodynamic equilibrium is a central challenge in statistical physics. Boltzmann Generators (BGs) tackle it by combining a generative model with a Monte Carlo (MC) correction step to obtain asymptotically unbiased samples from an unnormalized target. Most current BGs use classic MC mechanisms such as importance sampling, which both require tractable likelihoods from the backbone model and scale poorly in high-dimensional, multi-modal targets. We study BGs built on annealed Monte Carlo (aMC), which is designed to overcome these limitations by bridging a simple reference to the target through a sequence of intermediate densities. Diffusion models (DMs) are powerful generative models and have already been incorporated into aMC-based recalibration schemes via the diffusion-induced density path, making them appealing backbones for aMC-BGs. We provide an empirical meta-analysis of DM-based aMC-BGs on controlled multi-modal Gaussian mixtures (varying mode separation, number of modes, and dimension), explicitly disentangling inference effects from learning effects by comparing (i) a perfectly learned DM and (ii) a DM trained from data. Even with a perfect DM, standard integrations using only first-order stochastic denoising kernels fail systematically, whereas second-order denoising kernels can substantially improve performance when covariance information is available. We further propose a deterministic aMC integration based on first-order transport maps derived from DMs, which outperforms the stochastic first-order variant at higher computational cost. Finally, in the learned-DM setting, all DM-aMC variants struggle to produce accurate BGs; we trace the main bottleneck to inaccurate DM log-density estimation. Code available at https://github.com/h2o64/dabg.

Figures

Figures reproduced from arXiv: 2601.21026 by Louis Grenioux, Maxence Noble.

Figure 1
Figure 1. Figure 1: Sampling results for classic annealed samplers with diffusion (blue) and tempering (red) density paths, when targeting TwoModes (Top) and ManyModes (Bottom) distributions in idealized setting (A). For tempering paths, we display the best-performing result among all values of K. For diffusion paths, we display the results for all values of K : the darker the bar, the higher K. In particular, these configura… view at source ↗
Figure 2
Figure 2. Figure 2: Different diffusion-based swapping mechanisms for Replica Exchange. (Left) standard swap scheme, see Section 2.2, where samples are exchanged directly across noise levels without guidance, potentially moving into low-probability regions. (Middle) DM-based swaps using forward and backward Markov kernels, as proposed by Zhang et al. (2025b) (coined Diff-APT), see Section 3.2. (Right) DM￾based swaps using for… view at source ↗
Figure 3
Figure 3. Figure 3: DM-based aMC-BG results with annealed samplers using different mechanisms, when targeting TwoModes distribution in idealized setting (A) : (First row) Low distance, high dimension, (Second row) High distance, low dimension, (Third row) Middle distance, middle dimension, (Fourth row) Average running time of each sampler over the three TwoModes variants. Each group of bars with the same color corresponds to … view at source ↗
Figure 4
Figure 4. Figure 4: Realistic results of DM-based aMC-BG, when targeting TwoModes distribution with intermediate difficulty (middle mode distance, middle dimension) in setting (B) : (From top to bottom) the DM is trained via TSM+DSM, tSM+DSM, aLFPE+DSM or RNE+DSM objective with identical computational budget. Each group of bars with the same color corresponds to a specific aMC method, except for the last two groups on the rig… view at source ↗
Figure 5
Figure 5. Figure 5: Exact density paths bridging π base (last time index) to 1D Gaussian mixtures (first time index). (Left): the target is an instance of TwoModes defined as (3/4)N(−4, 0.5 2 ) + (1/4)N(+4, 1), (Right) the target is the 1D instance of ManyModes Noble et al. (2025) with 32 modes, (First and third columns) diffusion density path, (Second and fourth columns) tempering density path. Given a discretization of 128 … view at source ↗
Figure 6
Figure 6. Figure 6: Learned diffusion density paths bridging π base (last time index) to the same targets as in [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Sampling results via the total variation distance of the weight histograms for classic annealed samplers with diffusion (blue) and tempering (red) density paths, when targeting ManyModes. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p047_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Sampling results via sliced Wasserstein distance for classic annealed samplers with diffusion (blue) and tempering (red) density paths, when targeting TwoModes in different settings of mode spacing and dimension. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p048_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Sampling results via mode weight absolute error for classic annealed samplers with diffusion (blue) and tempering (red) density paths, when targeting TwoModes in different settings of mode spacing and dimension. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p048_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Diffusion-based aMC-BG results via sliced Wasserstein distance in idealized setting (A), when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p049_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Diffusion-based aMC-BG results via mode weight absolute error in idealized setting (A), when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p049_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Diffusion-based aMC-BG results via sliced Wasserstein distance in idealized setting (A), when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p050_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Diffusion-based aMC-BG results via mode weight absolute error in idealized setting (A), when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p050_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Diffusion-based aMC-BG results via sliced Wasserstein distance in idealized setting (A), when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p051_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Diffusion-based aMC-BG results via mode weight absolute error in idealized setting (A), when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p051_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Diffusion-based aMC-BG results via sliced Wasserstein distance in idealized setting (A), when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p052_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Diffusion-based aMC-BG results via mode weight absolute error in idealized setting (A), when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p052_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with DSM objective, when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p053_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with DSM objective, when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p053_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with DSM objective, when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p054_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with DSM objective, when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p054_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with TSM+DSM objective, when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p055_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with TSM+DSM objective, when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p055_23.png] view at source ↗
Figure 24
Figure 24. Figure 24: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with TSM+DSM objective, when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p056_24.png] view at source ↗
Figure 25
Figure 25. Figure 25: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with TSM+DSM objective, when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p056_25.png] view at source ↗
Figure 26
Figure 26. Figure 26: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with tSM+DSM objective, when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p057_26.png] view at source ↗
Figure 27
Figure 27. Figure 27: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with tSM+DSM objective, when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p057_27.png] view at source ↗
Figure 28
Figure 28. Figure 28: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with tSM+DSM objective, when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p058_28.png] view at source ↗
Figure 29
Figure 29. Figure 29: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with tSM+DSM objective, when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p058_29.png] view at source ↗
Figure 30
Figure 30. Figure 30: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with LFPE+DSM objective, when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p059_30.png] view at source ↗
Figure 31
Figure 31. Figure 31: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with LFPE+DSM objective, when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p059_31.png] view at source ↗
Figure 32
Figure 32. Figure 32: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with LFPE+DSM objective, when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p060_32.png] view at source ↗
Figure 33
Figure 33. Figure 33: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with LFPE+DSM objective, when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p060_33.png] view at source ↗
Figure 34
Figure 34. Figure 34: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with aLFPE+DSM objective, when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p061_34.png] view at source ↗
Figure 35
Figure 35. Figure 35: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with aLFPE+DSM objective, when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p061_35.png] view at source ↗
Figure 36
Figure 36. Figure 36: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with aLFPE+DSM objective, when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p062_36.png] view at source ↗
Figure 37
Figure 37. Figure 37: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with aLFPE+DSM objective, when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p062_37.png] view at source ↗
Figure 38
Figure 38. Figure 38: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with RNE+DSM objective, when targeting TwoModes distribution with a = 1, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p063_38.png] view at source ↗
Figure 39
Figure 39. Figure 39: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with RNE+DSM objective, when targeting TwoModes distribution with a = 2.5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p063_39.png] view at source ↗
Figure 40
Figure 40. Figure 40: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with RNE+DSM objective, when targeting TwoModes distribution with a = 5, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p064_40.png] view at source ↗
Figure 41
Figure 41. Figure 41: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with RNE+DSM objective, when targeting TwoModes distribution with a = 10, for all dimensional settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p064_41.png] view at source ↗
Figure 42
Figure 42. Figure 42: Diffusion-based aMC-BG results via sliced Wasserstein distance in idealized setting (A), when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p065_42.png] view at source ↗
Figure 43
Figure 43. Figure 43: Diffusion-based aMC-BG results via weight histogram total variation distance in idealized setting (A), when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p065_43.png] view at source ↗
Figure 44
Figure 44. Figure 44: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with DSM objective, when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p066_44.png] view at source ↗
Figure 45
Figure 45. Figure 45: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with TSM+DSM objective, when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p066_45.png] view at source ↗
Figure 46
Figure 46. Figure 46: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with tSM+DSM objective, when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p067_46.png] view at source ↗
Figure 47
Figure 47. Figure 47: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with LFPE+DSM objective, when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p067_47.png] view at source ↗
Figure 48
Figure 48. Figure 48: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with aLFPE+DSM objective, when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p068_48.png] view at source ↗
Figure 49
Figure 49. Figure 49: Diffusion-based aMC-BG results via sliced Wasserstein distance in realistic setting (B) with RNE+DSM objective, when targeting ManyModes distribution in all possible settings. This is complementary to [PITH_FULL_IMAGE:figures/full_fig_p068_49.png] view at source ↗
Figure 50
Figure 50. Figure 50: Diffusion density paths bridging π base (last time index) to the same TwoModes target as in [PITH_FULL_IMAGE:figures/full_fig_p069_50.png] view at source ↗
Figure 51
Figure 51. Figure 51: Diffusion density paths bridging π base (last time index) to the same ManyModes target as in [PITH_FULL_IMAGE:figures/full_fig_p070_51.png] view at source ↗
Figure 52
Figure 52. Figure 52: Sensitivity of SMC methods (including diffusion-based BGs) with respect to the number of local MCMC steps across different values of K. We do not consider any MCMC warm-up procedure here. The default setting of the main experiments is represented by a dashed line. D.4 Ablation studies In this section, we investigate the sensitivity of the annealed samplers considered in this work with respect to their res… view at source ↗
Figure 53
Figure 53. Figure 53: Sensitivity of RE methods (including diffusion-based BGs) with respect to the swap period across different values of K. The global number of MCMC steps is the same as in the main experiments. The default setting of the main experiments is represented by a dashed line. For the three best performing methods, we observe a sweet spot when varying the swap period. RE with base init. RE with score-informed init… view at source ↗
Figure 54
Figure 54. Figure 54: Sensitivity of diffusion-based RE-BGs methods with respect to the per-level initial￾ization across different values of K ∈ {16, 32, 64, 128, 256} : the darker the bar, the higher K. The group of bars displayed on the right of the figure (“RE with score-informed init”) exactly corresponds to the RE bars from the third row of [PITH_FULL_IMAGE:figures/full_fig_p072_54.png] view at source ↗
Figure 55
Figure 55. Figure 55: Sensitivity of all diffusion-based aMC-BGs with deterministic transitions with respect to the number of fixed-point iterations (x-axis) and the truncation order I, across different values of K. (Top) AIS variant, (Middle) SMC variant, (Bottom) RE variant. We fix NH = 32. 74 [PITH_FULL_IMAGE:figures/full_fig_p074_55.png] view at source ↗
Figure 56
Figure 56. Figure 56: Sensitivity of all diffusion-based aMC-BGs with deterministic transitions with respect to the number of Hutchinson samples NH (x-axis) and the truncation order I, across different values of K. (Top) AIS variant, (Middle) SMC variant, (Bottom) RE variant. We fix M = 4. 75 [PITH_FULL_IMAGE:figures/full_fig_p075_56.png] view at source ↗
Figure 57
Figure 57. Figure 57: Sensitivity of all diffusion-based aMC-BGs with deterministic transitions with respect to the number of Hutchinson samples NH (x-axis) and fixed-point iterations M, across different values of K. (Top) AIS variant, (Middle) SMC variant, (Bottom) RE variant. We fix I = 3. 76 [PITH_FULL_IMAGE:figures/full_fig_p076_57.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

56 extracted references · 4 canonical work pages

  1. [1]

    standardized

    + 1 3N(x;a1d,Σ 2), whereΣ 1, Σ2∈R d×d are diagonal covariance matrices. The diagonal entries ofΣ1 are given by(Σ1)i,i = i dσ2 max + d−i d σ2 min, and those ofΣ 2 are the reverse ofΣ 1:(Σ 2)i,i = (Σ 1)d−i,d−i, with σ2 max = 0.2and σ2 min = 0.01(hence, the conditioning number of each covariance matrix is20). We vary the separation parameter a∈{ 1.0, 2.5, 5....

  2. [2]

    Proof.This is an immediate corollary from (Hall, 2000, Theorem 3.6)

    For any matrixM∈R d×d satisfying∥M∥<min(1/|α|,1/|β|), the following identities hold log [ (Id−βM)−1(Id +αM) ] = ∞∑ i=1 βi−(−1)iαi i Mi,log [ (Id +βM)−1(Id−αM) ] = ∞∑ i=1 (−1)iβi−αi i Mi, wherelogdenotes the matrix logarithm. Proof.This is an immediate corollary from (Hall, 2000, Theorem 3.6). Corollary 5.Let( c1,c 2,c

  3. [3]

    A.2 Discrete time setting for diffusion models Following Karras et al

    Similar computations withM2 lead to the second result. A.2 Discrete time setting for diffusion models Following Karras et al. (2024); Grenioux et al. (2024), we define the time discretization{tk}K k=0⊂ [0,T ] accordingly to the growth (in log-scale) of the noise levelt↦→σ(t). 31 Given fixed boundary valuesσmin >0andσ max >0, we define for anyk∈{0,...,K}th...

  4. [6]

    =dlog|c 1|+ log det ( (Id +c 2M)−1(Id−c 3M ) ) =dlog|c 1|+ Tr log ( (Id +c 2M)−1(Id−c 3M ) ).(Hall, 2000, Theorem 3.10) Hence, we obtain the first result of Corollary 5 by using the second statement of Lemma 4 withβ =c2 and α=c

  5. [8]

    Gabriel Cardoso, Yazid Janati el idrissi, Sylvain Le Corff, and Eric Moulines

    URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ bcd11db0b26d8fc2266b91d3ff982ed1-Paper-Conference.pdf. Gabriel Cardoso, Yazid Janati el idrissi, Sylvain Le Corff, and Eric Moulines. Monte carlo guided denoising diffusion models for bayesian linear inverse problems. InThe Twelfth International Conference on Learning Representations,

  6. [9]

    Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud

    URL https://proceedings.neurips.cc/paper/2019/file/ 5d0d5594d24f0f955548f0fc0ff83d10-Paper.pdf. Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equations.Advances in neural information processing systems,

  7. [13]

    URLhttps://www.pnas.org/doi/abs/10.1073/pnas.2109420119

    doi: 10.1073/pnas.2109420119. URLhttps://www.pnas.org/doi/abs/10.1073/pnas.2109420119. Ruiqi Gao, Yang Song, Ben Poole, Ying Nian Wu, and Diederik P Kingma. Learning energy-based models by diffusion recovery likelihood. InInternational Conference on Learning Representations,

  8. [14]

    Florentin Guth, Zahra Kadkhodaie, and Eero P Simoncelli

    URL https: //openreview.net/forum?id=d91E9RhVFU. Florentin Guth, Zahra Kadkhodaie, and Eero P Simoncelli. Learning normalized image densities via dual score matching. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025a. URLhttps://openreview.net/forum?id=wtYcS4kxpF. Florentin Guth, Zahra Kadkhodaie, and Eero P Simoncelli. ...

  9. [16]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel

    URLhttps://arxiv.org/abs/2506.05668. Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851,

  10. [17]

    Koji Hukushima and Koji Nemoto

    URLhttps://arxiv.org/abs/2210.02303. Koji Hukushima and Koji Nemoto. Exchange monte carlo method and application to spin glass simulations. Journal of the Physical Society of Japan, 65(6):1604–1608,

  11. [19]

    D Experimental details D.1 Target details Definition of theTwoModestarget distribution.For our target π, we first consider the Gaussian mixture introduced in Grenioux et al

    Similarly, one could use the coefficients from Proposition 25 for the VE noising scheme combined with exponential integration. D Experimental details D.1 Target details Definition of theTwoModestarget distribution.For our target π, we first consider the Gaussian mixture introduced in Grenioux et al. (2025), whose density is defined overRd as γ(x) = 2 3N(x;−a1d,Σ

  12. [20]

    URLhttps://doi.org/10.1021/acs.jpclett.2c03327

    doi: 10.1021/acs.jpclett.2c03327. URLhttps://doi.org/10.1021/acs.jpclett.2c03327. Yazid Janati, Badr Moufad, Alain Durmus, Eric Moulines, and Jimmy Olsson. Divide-and-conquer posterior sampling for denoising diffusion priors. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.),Advances in Neural Information Processi...

  13. [21]

    Yazid Janati, Eric Moulines, Jimmy Olsson, and Alain Oliviero-Durmus

    URL https://proceedings.neurips.cc/paper_files/ paper/2024/file/b0ae046e198a5e43141519868a959c74-Paper-Conference.pdf. Yazid Janati, Eric Moulines, Jimmy Olsson, and Alain Oliviero-Durmus. Bridging diffusion posterior sampling and monte carlo methods: a survey.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Scienc...

  14. [22]

    URLhttps: //royalsocietypublishing.org/doi/abs/10.1098/rsta.2024.0331

    doi: 10.1098/rsta.2024.0331. URLhttps: //royalsocietypublishing.org/doi/abs/10.1098/rsta.2024.0331. Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models.Advances in Neural Information Processing Systems, 35:26565–26577,

  15. [23]

    Leon Klein, Andreas Krämer, and Frank Noé

    URLhttps://proceedings.neurips.cc/ paper_files/paper/2024/file/5035a409f5798e188079e236f437e522-Paper-Conference.pdf. Leon Klein, Andreas Krämer, and Frank Noé. Equivariant flow matching.Neural Information Processing Systems (NeurIPS),

  16. [26]

    doi: 10.1145/3341156

    ISSN 0730-0301, 1557-7368. doi: 10.1145/3341156. URLhttps://dl.acm.org/doi/10.1145/3341156. Radford M Neal. Annealed importance sampling.Statistics and computing, 11:125–139,

  17. [28]

    doi: 10.1126/science.aaw1147

    ISSN 0036-8075, 1095-9203. doi: 10.1126/science.aaw1147. URLhttps://www.science.org/doi/10.1126/science.aaw1147. Kaoru Ohno, Keivan Esfarjani, and Yoshiyuki Kawazoe.Computational Materials Science: From Ab Initio to Monte Carlo Methods. Springer,

  18. [29]

    doi: https://doi.org/10.1016/S0009-2614(01)00055-0

    ISSN 0009-2614. doi: https://doi.org/10.1016/S0009-2614(01)00055-0. URL https://www.sciencedirect. com/science/article/pii/S0009261401000550. Bernt Øksendal. Stochastic differential equations. InStochastic differential equations: an introduction with applications, pp. 38–50. Springer,

  19. [30]

    URLhttp://www.jstor.org/ stable/3318418

    ISSN 13507265. URLhttp://www.jstor.org/ stable/3318418. Tim Salimans and Jonathan Ho. Should EBMs model the energy or the score? InEnergy Based Models Workshop-ICLR 2021,

  20. [31]

    Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, and Rajesh Ranganath

    URL https: //arxiv.org/abs/2410.15336. Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, and Rajesh Ranganath. A general framework for inference-time scaling and steering of diffusion models. InForty- second International Conference on Machine Learning,

  21. [32]

    Jeffrey Mark Siskind

    URLhttps://openreview.net/forum?id= Jp988ELppQ. Jeffrey Mark Siskind. Automatic differentiation: Inverse accumulation mode. InProgram Transformations for ML Workshop at NeurIPS 2019,

  22. [33]

    org/abs/2101.03288

    URLhttps://arxiv. org/abs/2101.03288. Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. InThe Ninth International Conference on Learning Representations,

  23. [36]

    URL https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/rssb.12464

    doi: https://doi.org/10.1111/rssb.12464. URL https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/rssb.12464. Saifuddin Syed, Alexandre Bouchard-Côté, Kevin Chern, and Arnaud Doucet. Optimised annealed sequential monte carlo samplers,

  24. [37]

    James Thornton, Louis Béthune, Ruixiang ZHANG, Arwen Bradley, Preetum Nakkiran, and Shuangfei Zhai

    URLhttps://arxiv.org/abs/2408.12057. James Thornton, Louis Béthune, Ruixiang ZHANG, Arwen Bradley, Preetum Nakkiran, and Shuangfei Zhai. Controlled generation with distilled diffusion energy models and sequential monte carlo. InThe 28th International Conference on Artificial Intelligence and Statistics,

  25. [38]

    Pascal Vincent

    URLhttps://arxiv.org/ abs/2407.13734. Pascal Vincent. A connection between score matching and denoising autoencoders.Neural computation, 23 (7):1661–1674,

  26. [39]

    29 Luhuan Wu, Brian L

    URLhttps://arxiv.org/abs/2008.10087. 29 Luhuan Wu, Brian L. Trippe, Christian A Naesseth, John Patrick Cunningham, and David Blei. Practical and asymptotically exact conditional sampling in diffusion models. InThirty-seventh Conference on Neural Information Processing Systems,

  27. [40]

    Fengzhe Zhang, Jiajun He, Laurence I Midgley, Javier Antorán, and José Miguel Hernández-Lobato

    URL https://openreview.net/forum?id=Gn2izAiYzZ. Fengzhe Zhang, Jiajun He, Laurence I Midgley, Javier Antorán, and José Miguel Hernández-Lobato. Efficient and unbiased sampling of boltzmann distributions via consistency models.arXiv preprint arXiv:2409.07323,

  28. [41]

    Midgley, and José Miguel Hernández-Lobato

    Fengzhe Zhang, Laurence I. Midgley, and José Miguel Hernández-Lobato. Efficient and unbiased sampling from boltzmann distributions via variance-tuned diffusion models, 2025a. URLhttps://arxiv.org/abs/ 2505.21005. Leo Zhang, Peter Potaptchik, Jiajun He, Yuanqi Du, Arnaud Doucet, Francisco Vargas, Hai-Dang Dau, and Saifuddin Syed. Accelerated parallel tempe...

  29. [45]

    Define the matrices M1 =c 1(Id +c 2M)−1(Id−c 3M),M 2 =c−1 1 (Id−c 3M)−1(Id +c 2M)

    LetM ∈R d×d be a matrix satisfying ∥M∥< min(1/|c 2|,1/|c 3|). Define the matrices M1 =c 1(Id +c 2M)−1(Id−c 3M),M 2 =c−1 1 (Id−c 3M)−1(Id +c 2M). Then we have log|det M 1|=dlog|c 1|+ ∞∑ i=1 (−1)ici 2−ci 3 i Tr[Mi],log|det M 2|=−dlog|c 1|+ ∞∑ i=1 ci 3−(−1)ici 2 i Tr[Mi]. Proof. Consider such(c1,c 2,c 3)and such matrixM. Note that the assumption onc2 and c3 ...

  30. [48]

    DSM objective

    Then, for any pair of time-steps(s,t )such thatT≥t>s≥0, the ODE solutionY t givenY s =y s∈R d is defined by Yt = exp (∫t s f(u)du ) ys + ( exp (∫t s f(u)du ) −1 ) b. Proof.Let0≤s<t≤T, setZ t = exp(−F(t))Yt, whereF(t) = ∫t 0f(u)du, then dZt =f(t) exp(−F(t))bdt, which implies that Zt =Z+ (exp(−F(s))−exp(−F(t)))b, which gives the result. 32 A.4 Review of sco...

  31. [49]

    TSM objective

    2 2 ] , withX t =S(t)X 0 +r(t)Z , X− t =S(t)X 0−r(t)Z. In our experiments, the “TSM objective” will systematically refer to the training loss function˜Lanti TSM. As originally proposed by Bortoli et al. (2024), this loss can also be combined with preconditioning schemes to reduce its variance in practice; however, since those are not compatible with the...

  32. [50]

    Lemma 8(Exact noising SDE integration - General case).The conditional distribution of Xt given Xs =x s∈R d is defined by the Gaussian kernel qt|s(·|xs) = N ( {S(t)/S(s)}xs,S(t) 2{σ2(t)−σ 2(s)}Id ) , 9While He et al. (2025) propose to replaceqt|s, though tractable, by its Euler-Maruyama estimation, our implementation relies rather on its exact formulation ...

  33. [51]

    Below, we present a rigorous expression of this upper bound onδ for the noising schemes considered in this paper, that is theVariance-Preservingapproach (see Section B.2) and theVariance-Explodingapproach (see Section B.3). B.2 Variance-Preserving diffusion Consider the noising SDE(2) where f(t) =−g2(t)/2and g being such that ∫T 0 g2(s)ds≫ 1, with arbitra...

  34. [52]

    Then, SDE (2) simply writes as dXt =g(t)dW t, X0∼π .(31) This noising scheme is known as theVariance-Exploding(VE) scheme (Song et al., 2021). On the choice of theg-schedule.Following the guidelines from (Karras et al., 2022), we consider the geometric schedule g2(t) =σ 2 min (σ2 max σ2 min )t log (σ2 max σ2 min ) , whereσ min≈0andσ max≫1can be arbitraril...

  35. [55]

    To evaluate how well mode weights are recovered, we compute the Total Variation (TV) distance between the true mode weight histogram and its Monte Carlo estimate

    Moreover, we apply the same standardization procedure as for theTwoModestargets. To evaluate how well mode weights are recovered, we compute the Total Variation (TV) distance between the true mode weight histogram and its Monte Carlo estimate. D.2 Training and sampling parameters Diffusion model training details.As explained in Section 6.1, we consider tw...

  36. [56]

    In particular, when training networks with TSM and RNE objectives, we initializeUθ based on the output of DSM training procedure

    and set the default learning rate as10−4, multiplied by a factord−1 for score matching approaches (DSM, TSM, tSM) andd−2 for energy matching techniques (LFPE, aLFPE, RNE), following guidelines of related works. In particular, when training networks with TSM and RNE objectives, we initializeUθ based on the output of DSM training procedure. In the case of t...

  37. [1953]

    doi: 10.1063/1.1699114

    ISSN 0021-9606. doi: 10.1063/1.1699114. URLhttps://doi.org/10.1063/1.1699114. L. I. Midgley, V. Stimper, G. N. C. Simm, and J. M. Hernández-Lobato. Bootstrap your flow. In1st ELLIS Machine Learning for Molecule Discovery Workshop, December

  38. [1982]

    doi: https://doi.org/10.1016/0304-4149(82)90051-5

    ISSN 0304-4149. doi: https://doi.org/10.1016/0304-4149(82)90051-5. URLhttps: //www.sciencedirect.com/science/article/pii/0304414982900515. Christophe Andrieu and Gareth O. Roberts. The pseudo-marginal approach for efficient Monte Carlo computations.The Annals of Statistics, 37(2):697 – 725,

  39. [1986]

    URLhttps://link.aps.org/ doi/10.1103/PhysRevLett.57.2607

    doi: 10.1103/PhysRevLett.57.2607. URLhttps://link.aps.org/ doi/10.1103/PhysRevLett.57.2607. Publisher: American Physical Society. Saifuddin Syed, Vittorio Romaniello, Trevor Campbell, and Alexandre Bouchard-Cote. Parallel tempering on optimized paths. In Marina Meila and Tong Zhang (eds.),Proceedings of the 38th International Conference on Machine Learnin...

  40. [1987]

    doi: https://doi.org/10.1016/0370-2693(87)91197-X

    ISSN 0370-2693. doi: https://doi.org/10.1016/0370-2693(87)91197-X. URL https://www.sciencedirect.com/science/article/pii/037026938791197X. Daan Frenkel and Berend Smit.Understanding Molecular Simulation: from Algorithms to Applications. Elsevier,

  41. [1989]

    URLhttps://doi.org/10.1080/03610918908812806

    doi: 10.1080/ 03610918908812806. URLhttps://doi.org/10.1080/03610918908812806. Aapo Hyvärinen. Estimation of non-normalized statistical models by score matching.Journal of Machine Learning Research, 6(24):695–709,

  42. [1996]

    URL https://doi.org/10.1143/JPSJ.65.1604

    doi: 10.1143/JPSJ.65.1604. URL https://doi.org/10.1143/JPSJ.65.1604. M.F. Hutchinson. A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines.Communications in Statistics - Simulation and Computation, 18(3):1059–1076,

  43. [2000]

    Jiajun He, José Miguel Hernández-Lobato, Yuanqi Du, and Francisco Vargas

    URLhttps://arxiv.org/ abs/math-ph/0005032. Jiajun He, José Miguel Hernández-Lobato, Yuanqi Du, and Francisco Vargas. Rne: plug-and-play diffusion inference-time control and energy-based training,

  44. [2001]

    MCMC using hamiltonian dynamics.arXiv preprint arXiv:1206.1901,

    Radford M Neal. MCMC using hamiltonian dynamics.arXiv preprint arXiv:1206.1901,

  45. [2009]

    URLhttps: //doi.org/10.1214/07-AOS574

    doi: 10.1214/07-AOS574. URLhttps: //doi.org/10.1214/07-AOS574. Michael Arbel, Alex Matthews, and Arnaud Doucet. Annealed flow transport monte carlo. InInternational Conference on Machine Learning, pp. 318–330. PMLR,

  46. [2010]

    Dynamical measure transport and neural pde solvers for sampling.arXiv preprint arXiv:2407.07873, 2024a

    Jingtong Sun, Julius Berner, Lorenz Richter, Marius Zeinhofer, Johannes Müller, Kamyar Azizzadenesheli, and Anima Anandkumar. Dynamical measure transport and neural pde solvers for sampling.arXiv preprint arXiv:2407.07873, 2024a. Jingtong Sun, Julius Berner, Lorenz Richter, Marius Zeinhofer, Johannes Müller, Kamyar Azizzadenesheli, and Anima Anandkumar. D...

  47. [2011]

    doi: 10.1145/1944345.1944349

    ISSN 0004-5411. doi: 10.1145/1944345.1944349. URLhttps://doi.org/10.1145/1944345.1944349. Fan Bao, Chongxuan Li, Jiacheng Sun, Jun Zhu, and Bo Zhang. Estimating the optimal covariance with imperfect mean in diffusion probabilistic models. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.),Proceedings of t...

  48. [2016]

    URLhttps://doi.org/10.1080/10618600.2015.1060885

    doi: 10.1080/10618600.2015.1060885. URLhttps://doi.org/10.1080/10618600.2015.1060885. Yaxuan Zhu, Jianwen Xie, Ying Nian Wu, and Ruiqi Gao. Learning energy-based models by cooperative diffusion recovery likelihood. InThe Twelfth International Conference on Learning Representations,

  49. [2017]

    URLhttp://www.jstor.org/stable/26408299

    ISSN 08834237, 21688745. URLhttp://www.jstor.org/stable/26408299. Tara Akhound-Sadegh, Jarrid Rector-Brooks, Joey Bose, Sarthak Mittal, Pablo Lemos, Cheng-Hao Liu, Marcin Sendera, Siamak Ravanbakhsh, Gauthier Gidel, Yoshua Bengio, Nikolay Malkin, and Alexander Tong. Iterated denoising energy matching for sampling from boltzmann densities. InProceedings of...

  50. [2019]

    doi: 10.1103/PhysRevD.100.034515

    ISSN 2470-0010, 2470-0029. doi: 10.1103/PhysRevD.100.034515. URLhttps://link.aps.org/doi/10.1103/PhysRevD.100.034515. Brian D.O. Anderson. Reverse-time diffusion equation models.Stochastic Processes and their Applications, 12(3):313–326,

  51. [2020]

    doi: 10.1007/978-3-030-47845-2_8

    ISBN 978-3-030-47845-2. doi: 10.1007/978-3-030-47845-2_8. URL https://doi.org/10.1007/978-3-030-47845-2_8. Luigi Del Debbio, Joe Marsh Rossney, and Michael Wilson. Machine Learning Trivializing Maps: A First Step Towards Understanding How Flow-Based Samplers Scale Up.PoS, LATTICE2021:059,

  52. [2021]

    URLhttps://arxiv.org/abs/2111. 11510. Laurence Midgley, Vincent Stimper, Javier Antorán, Emile Mathieu, Bernhard Schölkopf, and José Miguel Hernández-Lobato. Se(3) equivariant augmented coupling flows. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.),Advances in Neural Information Processing Systems, volume 36, pp. 79200–79225...

  53. [2022]

    Pierre Del Moral, Arnaud Doucet, and Ajay Jasra

    doi: 10.22323/1.396.0059. Pierre Del Moral, Arnaud Doucet, and Ajay Jasra. Sequential monte carlo samplers.Journal of the Royal Statistical Society Series B: Statistical Methodology, 68(3):411–436,

  54. [2023]

    URLhttps://onlinelibrary.wiley.com/ doi/abs/10.1002/sta4.625

    doi: https://doi.org/10.1002/sta4.625. URLhttps://onlinelibrary.wiley.com/ doi/abs/10.1002/sta4.625. Yan Zhou, Adam M. Johansen, and John A.D. Aston. Toward automatic model comparison: An adaptive sequential monte carlo approach.Journal of Computational and Graphical Statistics, 25(3):701–726,

  55. [2024]

    doi: 10.1038/s41586-024-07487-w

    ISSN 1476-4687. doi: 10.1038/s41586-024-07487-w. URLhttps: //doi.org/10.1038/s41586-024-07487-w. S. Agapiou, O. Papaspiliopoulos, D. Sanz-Alonso, and A. M. Stuart. Importance sampling: Intrinsic dimension and computational cost.Statistical Science, 32(3):405–431,

  56. [2025]

    URLhttps: //arxiv.org/abs/2506.16471. M. S. Albergo, G. Kanwar, and P. E. Shanahan. Flow-based generative models for markov chain monte carlo in lattice field theory.Physical Review D, 100(3):034515, 8

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.