Pith. sign in

REVIEW 1 major objections 4 minor 1 cited by

Temperature-Annealed Boltzmann Generators

T0 review · 1 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Training a flow at 1200 K, then annealing it to 300 K by reweighting its own samples, captures all metastable peptide states at a fraction of the leading baseline's energy cost.

desk verdict A genuinely new two-phase training strategy that delivers the best hexapeptide sampling I've seen, but the written objective omits the internal-coordinate Jacobian and needs a fix before the math matches the code. read the letter →

arxiv 2501.19077 v2 pith:2F4WIFWB submitted 2025-01-31 cs.LG

classification cs.LG
keywords temperature-annealedBoltzmanngeneratorsnormalizingflowsreverseKullback-Leiblerdivergencemodecollapseimportancesamplingreweightingdistributionmetastablestatesalaninepeptides
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes temperature-annealed Boltzmann generators (TA-BG), a two-phase recipe for training a normalizing flow — an invertible neural transformation with an exactly computable density — to sample the Boltzmann distribution of a molecule, the equilibrium distribution of its conformations, without the mode collapse that usually afflicts energy-based training. In the first phase the flow is trained with the reverse Kullback-Leibler divergence at 1200 K, where free-energy barriers are low enough that the molecule's metastable conformations are connected and the mode-seeking loss covers all of them. In the second phase the flow is cooled to 300 K along a geometric ladder of about nine intermediate temperatures: at each rung the flow samples a buffer, every sample is reweighted to the next lower temperature, the buffer is resampled, and the flow is retrained on it with the mass-covering forward KLD. On alanine dipeptide, tetrapeptide, and hexapeptide, the method matches or beats the Flow Annealed Importance Sampling Bootstrap (FAB) baseline on nearly all metrics, while using roughly a third of the target-energy evaluations on the two smaller systems; on the hexapeptide the authors report that it is the only variational method tested that resolves all metastable states. If the claim holds, accurate equilibrium sampling of flexible molecules becomes substantially cheaper, which matters most when each energy evaluation is expensive.

What carries the argument

The load-bearing mechanism is a two-stage temperature protocol built on a normalizing flow. Stage one trains the flow with the reverse Kullback-Leibler divergence at an elevated temperature; the mode-seeking behaviour of that loss, which collapses the flow at 300 K, is neutralized at 1200 K because the barriers shrink and the high-probability regions interconnect, and a regularized energy function prevents diverging van der Waals terms from destabilizing training. Stage two cools the flow along a geometric temperature ladder $T_i = T_{\mathrm{start}}\left(T_{\mathrm{target}}/T_{\mathrm{start}}\right)^{(i-1)/(K-1)}$, where each rung resamples a buffer of the flow's own samples using importance weights for the next temperature and retrains with the forward, mass-covering KLD, which is what keeps the cooling phase collapse-free. The flow itself is built from monotonic rational-quadratic spline coupling layers acting on internal coordinates (bond lengths, angles, dihedrals), with circular splines so that the periodic dihedral angles keep their correct topology. A final fine-tuning iteration at the target temperature, with $T_{i+1}=T_i$, raises the effective sample size of the training buffer and improves the final metrics.

What would settle it

Run the TA-BG recipe on a molecule whose metastable basins remain separated by a free-energy barrier of several $k_BT$ even at 1200 K, with the true 300 K populations fixed by a long unbiased molecular-dynamics trajectory. Apply the paper's own starting-temperature protocol — the fraction of collapsed reverse-KLD runs as a function of $T_1$ — to that system: if the pretraining collapses at 1200 K, or the annealed 300 K distribution assigns near-zero weight to a basin that the trajectory visits, the claim that temperature annealing circumvents mode collapse fails for that regime. A quantitative companion check is the final negative log-likelihood and Ramachandran KLD on an independent test set compared with the FAB baseline at a matched target-energy budget.

Watch

Extended reading notes

Core claim

The central claim is that mode collapse in data-free normalizing-flow training is a temperature problem rather than an intrinsic failure of the reverse KLD. At 1200 K the free-energy barriers between metastable basins are low enough that reverse-KLD training covers all modes reliably, and the paper's starting-temperature ablation shows the fraction of collapsed runs dropping to zero there; at 300 K the same objective collapses, increasingly so for the larger peptides. The second claim is that an iterative reweighting anneal carries this coverage down to the target temperature: at each rung of a geometric temperature ladder, samples drawn from the flow at $T_i$ are weighted by $w(x)=p_{X,T_{i+1}}(x)/q_X(x;\theta)$, resampled, and used for forward-KLD training at $T_{i+1}$, repeated until 300 K. On the three alanine systems the resulting 300 K Ramachandran free-energy plots match long molecular-dynamics ground truth, with better negative log-likelihoods than FAB on all three systems and $7.56\times 10^7$ versus $2.13\times 10^8$ target energy evaluations on the two smaller systems. For the hexapeptide, the authors state, TA-BG is the only method tested that accurately resolves the metastable states.

Load-bearing premise

The load-bearing premise is that every new molecular system has some starting temperature at which reverse-KLD training reliably finds all metastable modes; the paper establishes this empirically for three alanine peptides, and if the pretraining misses a mode, the annealing steps can only reweight samples the already-collapsed flow covers and cannot recover it.

Editorial extensions

If this is right

  • Reverse-KLD training, previously discounted as unavoidably mode-collapsing, is sufficient for molecular Boltzmann sampling whenever the training temperature is high enough that metastable basins merge.
  • The buffered reweighting step is a collapse-free fine-tuning recipe for pretrained flows, applicable beyond temperature annealing, for example to debias flows trained on biased or non-equilibrated simulation data.
  • When target energies become the dominant cost, as with learned foundation-model force fields or ab initio potentials, the roughly threefold reduction in energy evaluations translates into a near-proportional wall-clock saving despite a larger number of flow evaluations.
  • Data-free variational sampling scales past the alanine-dipeptide benchmark: on the hexapeptide the only variational method that resolves all metastable states is TA-BG, making larger and more flexible molecules plausible targets.
  • Inside the same annealing ladder, plain importance sampling's exponential effective-sample-size decay in high dimensions can be replaced by annealed importance sampling, keeping buffer overlap approximately constant as system size grows.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An automatic pretraining protocol suggests itself: scan the starting temperature upward until the fraction of collapsed reverse-KLD runs drops to zero, following the paper's Figure 6 procedure, and then begin annealing; this converts the per-system empirical heuristic into a design rule for new molecules.
  • The critical temperature at which reverse-KLD training stops collapsing should track the physical barrier heights of the target's free-energy landscape, so a cheap low-temperature barrier estimate could predict a safe starting temperature without retraining.
  • The annealing buffer's effective sample size could serve as a live thermostat: pick the next temperature on the fly to hold the buffer overlap at a target value, adapting the schedule to systems whose modes merge faster or slower than alanine peptides.
  • Because the anneal is driven by forward KLD on reweighted samples rather than by the target energy directly, the annealing phase should preserve whatever modes the pretraining found; a testable consequence is that swapping the pretraining objective, for instance to FAB's $\alpha$-divergence, while keeping the anneal would leave the final mode coverage unchanged.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper proposes temperature-annealed Boltzmann generators (TA-BG), a variational sampling method that first trains a normalizing flow with the reverse Kullback-Leibler divergence at an elevated temperature (1200 K) to avoid mode collapse, and then anneals the learned distribution to a target temperature (300 K) through iterative importance-weighted resampling followed by forward-KLD training. The method is evaluated on alanine dipeptide, tetrapeptide, and hexapeptide in implicit solvent, where it is compared against flow models trained with forward KLD on MD data, reverse KLD at 300 K, and the FAB baseline. The authors report that TA-BG matches or improves on FAB on most metrics while using up to about three times fewer target energy evaluations, and that for the hexapeptide it is the only variational method that accurately resolves the metastable states. The paper also includes extensive ablations, a 2D Gaussian mixture study, and a public implementation and ground-truth datasets.

Significance. If the central claims hold, this is a practically useful contribution to data-free variational sampling of molecular systems. The main idea is simple and plausible: high-temperature reverse-KLD training avoids mode collapse, and a sequence of short annealing steps with reweighting transfers the learned density to the target temperature. The empirical evaluation is careful and unusually thorough: four independent runs per condition, standard errors, hyperparameter ablations for starting temperature, annealing schedule, buffer size, and fine-tuning, a robustness check for FAB, and a fair comparison of wall-clock time in Appendix J. The release of code and ground-truth data is a concrete asset. The main reservation is that the mathematical formulation of the target density in internal-coordinate space is incomplete, which affects the central claim of accurate Boltzmann sampling; this is a fixable issue but must be resolved before the paper can be accepted.

major comments (1)
  1. [Section 3.1 / 4.1, Eq. (6)] The training objective is written for a target density p_X defined on Cartesian coordinates, while the flow operates on internal coordinates (bond lengths, angles, dihedrals). The correct target density in internal coordinates is p_R(r) ∝ exp(−E(x_cart(r))/k_B T) |det(∂x_cart/∂r)|, and this Jacobian is not constant because bond lengths and angles are flexible. The Jacobian is never defined or mentioned in the paper. As written, both the reverse-KLD pretraining and the importance weights in Section 4.2 target a different density, so the central claim that the method samples the Boltzmann distribution is not supported by the equations presented. Please state explicitly what density the flow is trained to match, include the Jacobian factor in the target density if the internal-coordinate frame is used, and document how the released implementation handles this term.
minor comments (4)
  1. [Section 6] The sentence 'For FAB applied to the hexapeptide, even when using almost 3 times as many target evaluations compared to our approach, we still achieve a lower NLL value' is ambiguous; it should be rephrased to make clear that TA-BG achieves the lower NLL, while FAB uses more evaluations.
  2. [Appendix F.3, Figure 6] Mode collapse in the starting-temperature ablation is defined by 'manual visual inspection of the Ramachandran plots'; please specify a more objective criterion or at least note the possible subjectivity of this threshold.
  3. [Section 4.2] The statement 'mode collapse is not a problem during the annealing' is too strong: forward-KLD training can only fit the support represented by reweighted samples from the current flow, so a mode missed by the high-temperature pretraining cannot be recovered. The limitation is acknowledged in Appendix F.3, but the main text should state this caveat.
  4. [Appendix J, Table 14] The wall-time comparison is useful, but the main-text efficiency claim ('up to three times fewer target energy evaluations') should be explicitly distinguished from total wall-clock time, since the appendix shows that FAB currently has lower wall time on these systems.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; minor self-citations for architecture and training heuristics are not load-bearing, and the central claim is validated against independent MD ground truth and external baselines.

full rationale

The paper's central claim—that reverse-KLD pretraining at 1200 K followed by iterative reweighting-based annealing yields accurate 300 K Boltzmann distributions—is not circular. The target distribution enters only through force-field energy evaluations (Eq. 6 and the importance weights in Section 4.2); no parameter is fitted to the ground-truth MD samples used for evaluation. The NLL, RAM KLD, and ESS metrics are computed against independently simulated MD datasets (Section D), and the main baseline FAB is an external method (Midgley et al., 2023b), so the comparison does not reduce to the paper's own assumptions. The only self-citations (Schopmans & Friederich, 2024) are for the internal-coordinate representation, the spline architecture, and the heuristic of clipping the largest per-batch energy values; these are auxiliary design choices also attributed to external work and do not determine the central result. The paper also explicitly credits replica-exchange MD for the high-temperature idea, so it is not renaming a known result. A separate correctness concern exists: Eq. 6 is written for Cartesian coordinates while the flow operates on internal coordinates, and the manuscript does not state the corresponding Jacobian factor; however, this is a potential implementation/derivation error, not a circular reduction of the prediction to its inputs. Overall, the derivation chain is self-contained; the minor self-citations warrant score 2 rather than 0.

Assumptions & free parameters 11 free parameters · 6 assumptions · 0 invented entities

The central claim rests on the Boltzmann target defined by an AMBER force field, the expressiveness of the chosen normalizing flow, and the validity of importance-sampling-based annealing. No new physical entities are introduced. The method relies on a number of empirically tuned hyperparameters, listed above, which are not derived from theory.

free parameters (11)
  • Starting temperature T1 = 1200 K
    Chosen above the empirically determined critical temperature for mode collapse (Figure 6); system-specific and not predicted from theory.
  • Number of annealing iterations K = 9 (plus final fine-tuning)
    Ablation shows more iterations improve NLL and ESS with diminishing returns (Table 8); choice balances accuracy vs target evaluations.
  • Temperature schedule ratio = Geometric progression over 9 steps
    Geometric schedule keeps buffer ESS approximately constant compared to linear (Figures 7-8).
  • Buffer sample sizes = 5e6 (di/tetra) or 1e7 (hexa) drawn, resampled to 2e6
    Ablation on dipeptide shows increasing buffer size improves metrics up to a point (Table 9).
  • Gradient steps per annealing iteration = 30,000 (dipeptide), 20,000 (tetra/hexa)
    Hyperparameter chosen by hand; larger systems use fewer steps.
  • Learning rate = 5e-6 or 1e-5 depending on system
    Empirical choice; no sensitivity study reported.
  • Importance weight clipping threshold = 0.01% highest weights clipped
    Prevents outliers from dominating training and metrics (Sections E and F.2).
  • Energy regularization parameters = Ehigh=1e8, Emax=1e20
    Taken from FAB; stabilizes training against atom clashes (Equation 10).
  • Number of highest-energy values removed per batch = 10 (dipeptide at 1200K), 20 (hexapeptide), 40 (300K runs)
    Heuristic to stabilize reverse KLD training (Section H.1).
  • Intermediate fine-tuning (hexapeptide) = One extra iteration after each annealing step
    Ablation shows it is needed only for the hexapeptide to avoid severe ESS drop (Table 7, Figure 9).
  • Scaling constants for internal coordinates = sigma=0.07 nm for bonds, 0.5730 rad for angles
    Empirically chosen to map internal coordinates into the spline range [0,1] (Equation 9).
assumptions (6)
  • domain assumption The AMBER force field energy E(x) defines the correct Boltzmann target p(x) proportional to exp(-E/kBT).
    All training and ground-truth data use AMBER with implicit solvent; the flow is evaluated against MD samples from the same model, so this is a self-consistent but not externally validated target.
  • domain assumption The normalizing flow architecture (16 neural spline coupling layers) is expressive enough to represent the Boltzmann distributions at all temperatures from 1200K to 300K.
    No universal approximation guarantee for this restricted architecture; empirical success on three systems suggests adequacy.
  • domain assumption The internal coordinate representation (bonds, angles, dihedrals) with fixed scalings is a sufficient coordinate system for the conformational distribution.
    Removes translation and rotation symmetry but ignores permutation symmetry of identical atoms, acknowledged as a limitation in Section 7.
  • standard math Self-normalized importance sampling produces unbiased estimates of target expectations given enough samples.
    Standard importance sampling theory (Equation 7), but practical accuracy degrades when overlap is low.
  • domain assumption Reweighted forward KLD training on resampled datasets drives the flow toward the target distribution at the next temperature.
    This is maximum likelihood on approximately distributed samples; convergence is not proven and relies on sufficient buffer ESS.
  • domain assumption The flow pre-trained at 1200K does not need reinitialization during annealing.
    Empirically, continuous parameter updates work; no theory prevents catastrophic forgetting between iterations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Temperature-Annealed Boltzmann Generators." pith.science (2026). https://pith.science/paper/2F4WIFWB

@misc{pith2026250119077,
  author       = {Pith},
  title        = {Pith review of: Temperature-Annealed Boltzmann Generators},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2F4WIFWB}},
  note         = {Machine review of arXiv:2501.19077}
}
read the original abstract

Efficient sampling of unnormalized probability densities such as the Boltzmann distribution of molecular systems is a longstanding challenge. Next to conventional approaches like molecular dynamics or Markov chain Monte Carlo, variational approaches, such as training normalizing flows with the reverse Kullback-Leibler divergence, have been introduced. However, such methods are prone to mode collapse and often do not learn to sample the full configurational space. Here, we present temperature-annealed Boltzmann generators (TA-BG) to address this challenge. First, we demonstrate that training a normalizing flow with the reverse Kullback-Leibler divergence at high temperatures is possible without mode collapse. Furthermore, we introduce a reweighting-based training objective to anneal the distribution to lower target temperatures. We apply this methodology to three molecular systems of increasing complexity and, compared to the baseline, achieve better results in almost all metrics while requiring up to three times fewer target energy evaluations. For the largest system, our approach is the only method that accurately resolves the metastable states of the system.

Figures

Figures reproduced from arXiv: 2501.19077 by the authors.

Figure 1
Figure 1. (a) Illustration of Boltzmann generators based on normalizing flows. The goal is to learn the equilibrium Boltzmann distribution of the 3D conformations of a molecular system. We focus on data-free training, where only the unnormalized probability density is known. (b) Illustration of our workflow. (1) To avoid mode collapse, we first train the flow at high temperature with the reverse KLD. (2-4) Then, the learned d… view at source ↗
Figure 2
Figure 2. Visualization of the iterative annealing process, showing the free energy F = −kBT ln p(ϕi, ψi) of backbone dihedral angles (Ramachandran plots) in each iteration. After learning the distribution at 1200 K using the reverse KLD, the distribution is annealed step by step to the target temperature 300 K. Note that not all annealing iterations are shown. Since the tetrapeptide has three pairs of backbone dihedral angle… view at source ↗
Figure 3
Figure 3. Comparison of the free energy F = −kBT ln p(ϕi, ψi) of the backbone dihedral angles (Ramachandran plots) at 300 K obtained by the different methods. Since the tetrapeptide has three pairs of backbone dihedral angles, and the hexapeptide five, we selected the pair with the largest visible deviation among the methods. The Ramachandran plots for all pairs of backbone dihedral angles can be found in the appendix, Figure… view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Illustration of a normalizing flow coupling layer. For the normalizing flow architecture, we use an architecture similar to previous works (Midgley et al., 2023b; Schopmans & Friederich, 2024). As the invertible transformation in the coupling layers, we use monotonic r…
Figure 5
Figure 5. Figure 5: Free energy F = −kBT ln p(ϕi, ψi) of dihedral angles (Ramachandran plots), reweighted directly from 1200 K to 300 K. We used 1 × 106 samples for importance sampling (middle). T1 = 1200 K. A tradeoff exists: Increasing T1 allows cheaper reverse KLD training without mode…
Figure 6
Figure 6. Figure 6: Ablation of starting temperature T1: Fraction of reverse KLD experiments with mode collapse at a given temperature. 1200 1000 800 600 400 Ti / K 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Buffer ESS Buffer ESSTi ! Ti+1 (linear temp schedule) Buffer ESSTi ! Ti+1 (geometric temp sc…
Figure 7
Figure 7. Figure 7: The ESS of the training buffer datasets W used to anneal from Ti to Ti+1, comparing a linear and geometric temperature schedule for alanine dipeptide. The increase of the buffer ESS in the end is due to the final fine-tuning iteration with Ti+1 = Ti. 19 [PITH_FULL_IMA…
Figure 8
Figure 8. Figure 8: The ESS of the training buffer datasets W used to anneal from Ti to Ti+1, comparing a linear and geometric temperature schedule for alanine hexapeptide. Note that the hexapeptide system includes intermediate fine-tuning iterations (Ti+1 = Ti), where the buffer ESS is i…
Figure 9
Figure 9. Figure 9: The ESS of the training buffer datasets W used to anneal from Ti to Ti+1, comparing the case with and without intermediate fine-tuning iterations for the hexapeptide system [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10 [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: Comparison of obtained distributions for the 2D GMM system. We visualize 1000 samples from the ground truth distribution, TA-BG, and FAB. We additionally visualize the contour lines of the ground truth log probability. We report the obtained NLL and ESS in [PITH_FULL…
Figure 12
Figure 12. Figure 12: Reweighted version of [PITH_FULL_IMAGE:figures/full_fig_p028_12.png]
Figure 13
Figure 13. Figure 13: Visualization of the iterative annealing process for the tetrapeptide, showing the free energy F = −kBT ln p(ϕi, ψi) of backbone dihedral angles (Ramachandran plots) in each iteration. Note that not all annealing iterations are shown. We used 1 × 107 samples for the R…
Figure 14
Figure 14. Figure 14: Comparison of the free energy F = −kBT ln p(ϕi, ψi) of the backbone dihedral angles (Ramachandran plots) of the tetrapeptide at 300 K. 30 [PITH_FULL_IMAGE:figures/full_fig_p030_14.png]
Figure 15
Figure 15. Figure 15: Reweighted version of [PITH_FULL_IMAGE:figures/full_fig_p031_15.png]
Figure 16
Figure 16. Figure 16: Visualization of the iterative annealing process for the hexapeptide, showing the free energy F = −kBT ln p(ϕi, ψi) of backbone dihedral angles (Ramachandran plots) in each iteration. Note that not all annealing iterations are shown. We used 1 × 107 samples for the Ra…
Figure 17
Figure 17. Figure 17: Comparison of the free energy F = −kBT ln p(ϕi, ψi) of the backbone dihedral angles (Ramachandran plots) of the hexapeptide at 300 K. 33 [PITH_FULL_IMAGE:figures/full_fig_p033_17.png]
Figure 18
Figure 18. Figure 18: Reweighted version of [PITH_FULL_IMAGE:figures/full_fig_p034_18.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework

    cs.HC 2025-08 unverdicted novelty 5.0 of 10

    A three-layer framework (input, processing, output) for adaptive external human-machine interfaces in autonomous vehicles is introduced to systematize design and analysis.

Reference graph

Works this paper leans on

53 extracted references · 35 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    J., Bambrick, J., Bodenstein, S

    Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A. J., Bambrick, J., Bodenstein, S. W., Evans, D. A., Hung, C.-C., O'Neill, M., Reiman, D., Tunyasuvunakool, K., Wu, Z., Z emgulyt \.e , A., Arvaniti, E., Beattie, C., Bertolli, O., Bridgland, A., Cherepanov, A., Congreve, M., Cowen-Rivers , A. ...

  3. [3]

    Iterated Denoising Energy Matching for Sampling from Boltzmann Densities

    Akhound-Sadegh , T., Rector-Brooks , J., Bose, J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., Malkin, N., and Tong, A. Iterated Denoising Energy Matching for Sampling from Boltzmann Densities . In Forty-First International Conference on Machine Learning , June 2024

  4. [4]

    Metadynamics

    Barducci, A., Bonomi, M., and Parrinello, M. Metadynamics. WIREs Computational Molecular Science, 1 0 (5): 0 826--843, 2011. ISSN 1759-0884. doi:10.1002/wcms.31

  5. [5]

    An optimal control perspective on diffusion-based generative modeling

    Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling. Transactions on Machine Learning Research, October 2023. ISSN 2835-8856

  6. [6]

    Beyond ELBOs : A Large-Scale Evaluation of Variational Methods for Sampling

    Blessing, D., Jia, X., Esslinger, J., Vargas, F., and Neumann, G. Beyond ELBOs : A Large-Scale Evaluation of Variational Methods for Sampling . In Forty-First International Conference on Machine Learning , June 2024

  7. [7]

    Deep learning the slow modes for rare events sampling

    Bonati, L., Piccini, G., and Parrinello, M. Deep learning the slow modes for rare events sampling. Proceedings of the National Academy of Sciences, 118 0 (44): 0 e2113533118, November 2021. doi:10.1073/pnas.2113533118

  8. [8]

    K., Bhikadiya, C., Bi, C., Bittrich, S., Chen, L., Crichlow, G

    Burley, S. K., Bhikadiya, C., Bi, C., Bittrich, S., Chen, L., Crichlow, G. V., Christie, C. H., Dalenberg, K., Di Costanzo, L., Duarte, J. M., Dutta, S., Feng, Z., Ganesan, S., Goodsell, D. S., Ghosh, S., Green, R. K., Guranovi \'c , V., Guzenko, D., Hudson, B. P., Lawson, C. L., Liang, Y., Lowe, R., Namkoong, H., Peisach, E., Persikova, I., Randle, C., R...

Show all 53 references
  1. [9]

    Nflows: Normalizing flows in PyTorch , November 2020

    Conor Durkan , Artur Bekasov , Iain Murray , and George Papamakarios . Nflows: Normalizing flows in PyTorch , November 2020

  2. [10]

    Case , H.M

    D.A. Case , H.M. Aktulga , K. Belfon , I.Y. Ben-Shalom , J.T. Berryman , S.R. Brozell , D.S. Cerutti , T.E. Cheatham, III , V.W.D. Cruzeiro , T.A. Darden , N. Forouzesh , G. Giambasu , T. Giese , M.K. Gilson , H. Gohlke , A.W. Goetz , J. Harris , S. Izadi , S.A. Izmailov , K. ...

  3. [11]

    Temperature steerable flows and Boltzmann generators

    Dibak, M., Klein, L., Kr \"a mer, A., and No \'e , F. Temperature steerable flows and Boltzmann generators. Phys. Rev. Res., 4 0 (4): 0 L042005, October 2022. doi:10.1103/PhysRevResearch.4.L042005

  4. [12]

    NICE : Non-linear Independent Components Estimation , April 2015

    Dinh, L., Krueger, D., and Bengio, Y. NICE : Non-linear Independent Components Estimation , April 2015

  5. [13]

    Density estimation using Real NVP , February 2017

    Dinh, L., Sohl-Dickstein , J., and Bengio, S. Density estimation using Real NVP , February 2017

  6. [14]

    On the Universality of Volume-Preserving and Coupling-Based Normalizing Flows

    Draxler, F., Wahl, S., Schnoerr, C., and Koethe, U. On the Universality of Volume-Preserving and Coupling-Based Normalizing Flows . In Forty-First International Conference on Machine Learning , June 2024

  7. [15]

    D., Pendleton, B

    Duane, S., Kennedy, A. D., Pendleton, B. J., and Roweth, D. Hybrid Monte Carlo . Physics Letters B, 195 0 (2): 0 216--222, September 1987. ISSN 0370-2693. doi:10.1016/0370-2693(87)91197-X

  8. [16]

    Neural Spline Flows

    Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G. Neural Spline Flows . In Advances in Neural Information Processing Systems , volume 32. Curran Associates, Inc., 2019

  9. [17]

    P., Abreu, C

    Eastman, P., Galvelis, R., Pel \'a ez, R. P., Abreu, C. R. A., Farr, S. E., Gallicchio, E., Gorenko, A., Henry, M. M., Hu, F., Huang, J., Kr \"a mer, A., Michel, J., Mitchell, J. A., Pande, V. S., Rodrigues, J. P., Rodriguez-Guerra , J., Simmonett, A. C., Singh, S., Swails, J....

  10. [18]

    Designing losses for data-free training of normalizing flows on Boltzmann distributions, January 2023

    Felardos, L., H \'e nin, J., and Charpiat, G. Designing losses for data-free training of normalizing flows on Boltzmann distributions, January 2023

  11. [19]

    Adaptive Annealed Importance Sampling with Constant Rate Progress

    Goshtasbpour, S., Cohen, V., and Perez-Cruz , F. Adaptive Annealed Importance Sampling with Constant Rate Progress . In Proceedings of the 40th International Conference on Machine Learning , pp.\ 11642--11658. PMLR, July 2023

  12. [20]

    Stochastic Optimal Control for Collective Variable Free Sampling of Molecular Transition Paths

    Holdijk, L., Du, Y., Hooft, F., Jaini, P., Ensing, B., and Welling, M. Stochastic Optimal Control for Collective Variable Free Sampling of Molecular Transition Paths . Advances in Neural Information Processing Systems, 36: 0 79540--79556, December 2023

  13. [21]

    Skipping the Replica Exchange Ladder with Normalizing Flows

    Invernizzi, M., Kr \"a mer, A., Clementi, C., and No \'e , F. Skipping the Replica Exchange Ladder with Normalizing Flows . J. Phys. Chem. Lett., 13 0 (50): 0 11643--11649, December 2022. doi:10.1021/acs.jpclett.2c03327

  14. [22]

    Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Z \'i dek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes , B., Nikolov, S., Jain, R., Adler, J., Back, T., Pet...

  15. [23]

    Kingma, D. P. and Ba, J. Adam: A Method for Stochastic Optimization , January 2017

  16. [24]

    and Noe, F

    Klein, L. and Noe, F. Transferable Boltzmann Generators . In The Thirty-eighth Annual Conference on Neural Information Processing Systems , November 2024

  17. [25]

    Lewis, S., Hempel, T., Jim \'e nez-Luna , J., Gastegger, M., Xie, Y., Foong, A. Y. K., Satorras, V. G., Abdin, O., Veeling, B. S., Zaporozhets, I., Chen, Y., Yang, S., Schneuing, A., Nigam, J., Barbero, F., Stimper, V., Campbell, A., Yim, J., Lienen, M., Shi, Y., Zheng, S., Sc...

  18. [26]

    H., Masters, M., Lee, S

    Mahmoud, A. H., Masters, M., Lee, S. J., and Lill, M. A. Accurate Sampling of Macromolecular Conformations Using Adaptive Deep Learning and Coarse-Grained Representation . J. Chem. Inf. Model., 62 0 (7): 0 1602--1617, April 2022. ISSN 1549-9596. doi:10.1021/acs.jcim.1c01438

  19. [27]

    Effective Sample Size for Importance Sampling based on discrepancy measures

    Martino, L., Elvira, V., and Louzada, F. Effective Sample Size for Importance Sampling based on discrepancy measures. Signal Processing, 131: 0 386--401, February 2017. ISSN 01651684. doi:10.1016/j.sigpro.2016.08.025

  20. [28]

    J., and Doucet, A

    Matthews, A., Arbel, M., Rezende, D. J., and Doucet, A. Continual Repeated Annealed Flow Transport Monte Carlo . In Proceedings of the 39th International Conference on Machine Learning , pp.\ 15196--15219. PMLR, June 2022

  21. [29]

    I., Stimper, V., Antor \'a n, J., Mathieu, E., Sch \"o lkopf, B., and Hern \'a ndez-Lobato , J

    Midgley, L. I., Stimper, V., Antor \'a n, J., Mathieu, E., Sch \"o lkopf, B., and Hern \'a ndez-Lobato , J. M. SE (3) Equivariant Augmented Coupling Flows . In Thirty-Seventh Conference on Neural Information Processing Systems , August 2023 a

  22. [30]

    I., Stimper, V., Simm, G

    Midgley, L. I., Stimper, V., Simm, G. N. C., Sch \"o lkopf, B., and Hern \'a ndez-Lobato , J. M. Flow Annealed Importance Sampling Bootstrap . In The Eleventh International Conference on Learning Representations , 2023 b . doi:10.48550/arXiv.2208.01893

  23. [31]

    Neal, R. M. Annealed importance sampling. Statistics and Computing, 11 0 (2): 0 125--139, April 2001. ISSN 1573-1375. doi:10.1023/A:1008923215028

  24. [32]

    No \'e , F. Bgflow. AI4Science group, FU Berlin (Frank No \'e and co-workers), December 2024

  25. [33]

    Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning

    No \'e , F., Olsson, S., K \"o hler, J., and Wu, H. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science, 365 0 (6457): 0 eaaw1147, September 2019. doi:10.1126/science.aaw1147

  26. [34]

    PyTorch : An Imperative Style , High-Performance Deep Learning Library , December 2019

    Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., K \"o pf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch : An I...

  27. [35]

    Graph neural networks for materials science and chemistry

    Reiser, P., Neubert, M., Eberhard, A., Torresi, L., Zhou, C., Shao, C., Metni, H., van Hoesel , C., Schopmans, H., Sommer, T., and Friederich, P. Graph neural networks for materials science and chemistry. Commun Mater, 3 0 (1): 0 1--18, November 2022. ISSN 2662-4443. doi:10.10...

  28. [36]

    J., Papamakarios, G., Racaniere, S., Albergo, M., Kanwar, G., Shanahan, P., and Cranmer, K

    Rezende, D. J., Papamakarios, G., Racaniere, S., Albergo, M., Kanwar, G., Shanahan, P., and Cranmer, K. Normalizing Flows on Tori and Spheres . In Proceedings of the 37th International Conference on Machine Learning , pp.\ 8083--8092. PMLR, November 2020

  29. [37]

    and Berner, J

    Richter, L. and Berner, J. Improved sampling via learned diffusions. In The Twelfth International Conference on Learning Representations , October 2023

  30. [38]

    Efficient mapping of phase diagrams with conditional Boltzmann Generators

    Schebek, M., Invernizzi, M., No \'e , F., and Rogal, J. Efficient mapping of phase diagrams with conditional Boltzmann Generators . Mach. Learn.: Sci. Technol., 5 0 (4): 0 045045, November 2024. ISSN 2632-2153. doi:10.1088/2632-2153/ad849d

  31. [39]

    and Friederich, P

    Schopmans, H. and Friederich, P. Conditional Normalizing Flows for Active Learning of Coarse-Grained Molecular Representations . In Forty-First International Conference on Machine Learning , June 2024

  32. [40]

    Improved off-policy training of diffusion samplers

    Sendera, M., Kim, M., Mittal, S., Lemos, P., Scimeca, L., Rector-Brooks , J., Adam, A., Bengio, Y., and Malkin, N. Improved off-policy training of diffusion samplers. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , November 2024

  33. [41]

    Y., and Ahn, S

    Seong, K., Park, S., Kim, S., Kim, W. Y., and Ahn, S. Transition Path Sampling with Improved Off-Policy Training of Diffusion Path Samplers . In The Thirteenth International Conference on Learning Representations , October 2024

  34. [42]

    A theoretical perspective on mode collapse in variational inference, October 2024

    Soletskyi, R., Gabri \'e , M., and Loureiro, B. A theoretical perspective on mode collapse in variational inference, October 2024

  35. [43]

    I., Simm, G

    Stimper, V., Midgley, L. I., Simm, G. N. C., Sch \"o lkopf, B., and Hern \'a ndez-Lobato , J. M. Alanine dipeptide in an implicit solvent at 300K . August 2022. doi:10.5281/zenodo.6993124

  36. [44]

    and Okamoto, Y

    Sugita, Y. and Okamoto, Y. Replica-exchange molecular dynamics method for protein folding. Chemical Physics Letters, 314 0 (1): 0 141--151, November 1999. ISSN 0009-2614. doi:10.1016/S0009-2614(99)01123-9

  37. [45]

    B., Bose, A

    Tan, C. B., Bose, A. J., Lin, C., Klein, L., Bronstein, M. M., and Tong, A. Scalable Equilibrium Sampling with Sequential Boltzmann Generators , February 2025

  38. [46]

    S., and Doucet, A

    Vargas, F., Grathwohl, W. S., and Doucet, A. Denoising Diffusion Samplers . In The Eleventh International Conference on Learning Representations , September 2022

  39. [47]

    Transport meets Variational Inference : Controlled Monte Carlo Diffusions

    Vargas, F., Padhy, S., Blessing, D., and N \"u sken, N. Transport meets Variational Inference : Controlled Monte Carlo Diffusions . In The Twelfth International Conference on Learning Representations , October 2023

  40. [48]

    TRADE : Transfer of Distributions between External Conditions with Normalizing Flows

    Wahl, S., Rousselot, A., Draxler, F., and Koethe, U. TRADE : Transfer of Distributions between External Conditions with Normalizing Flows . In The 28th International Conference on Artificial Intelligence and Statistics , February 2025

  41. [49]

    and Ahn, S

    Woo, D. and Ahn, S. Iterated Energy-based Flow Matching for Sampling from Boltzmann Densities , August 2024

  42. [50]

    A., Jaitly, N., and Susskind, J

    Zhai, S., Zhang, R., Nakkiran, P., Berthelot, D., Gu, J., Zheng, H., Chen, T., Bautista, M. A., Jaitly, N., and Susskind, J. Normalizing Flows are Capable Generative Models , December 2024

  43. [51]

    Zhang, D., Chen, R. T. Q., Liu, C.-H., Courville, A., and Bengio, Y. Diffusion Generative Flow Samplers : Improving learning signals through partial trajectory optimization. In The Twelfth International Conference on Learning Representations , October 2023

  44. [52]

    and Chen, Y

    Zhang, Q. and Chen, Y. Path Integral Sampler : A Stochastic Control Approach For Sampling . In International Conference on Learning Representations , October 2021

  45. [53]

    Predicting equilibrium distributions for molecular systems with deep learning

    Zheng, S., He, J., Liu, C., Shi, Y., Lu, Z., Feng, W., Ju, F., Wang, J., Zhu, J., Min, Y., Zhang, H., Tang, S., Hao, H., Jin, P., Chen, C., No \'e , F., Liu, H., and Liu, T.-Y. Predicting equilibrium distributions for molecular systems with deep learning. Nat Mach Intell, 6 0 ...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.