Pith. sign in

REVIEW 3 major objections 6 minor 69 references

Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

T0 review · 3 major / 6 minor · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Parallel Trajectory Tempering keeps energy-based models in equilibrium during training at roughly the cost of standard methods, and yields better samples on hard scientific data.

desk verdict Practical reservoir+adaptive PTT turns equilibrium RBM training into something that actually works on hard multimodal scientific data; quality evidence is solid, cost-parity is asserted not measured. read the letter →

arxiv 2607.27077 v1 pith:2UW6AF4V submitted 2026-07-29 cs.LG cond-mat.dis-nncond-mat.stat-mech

classification cs.LGcond-mat.dis-nncond-mat.stat-mech
keywords energy-basedmodelsrestrictedBoltzmannmachinesparalleltrajectorytemperingequilibriumsamplingpersistentcontrastivedivergencescientificgenerativemodelinglog-likelihoodestimationmultimodaldata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Energy-based models give an interpretable Boltzmann picture of data, but they are hard to train because Markov chains stop mixing as the energy landscape gets structured. This paper shows that Parallel Trajectory Tempering—exchanging samples between neighboring models along the learning path, not across temperatures—can keep sampling at equilibrium throughout training. With a reservoir of frozen equilibrium samples and adaptive learning-rate control, the cost stays comparable to Persistent Contrastive Divergence. On Restricted Boltzmann Machines the method beats standard EBM trainers and, on discrete tabular and scientific datasets, also beats stronger deep generative baselines in sample quality and resistance to overfitting when data are scarce or multimodal. As a side product it supplies thermalization times, equilibrium samples, and accurate log-likelihoods essentially for free.

What carries the argument

Parallel Trajectory Tempering (PTT): a ladder of model checkpoints frozen along the training trajectory; replicas swap between neighbors by a Metropolis rule on energy differences, while a reservoir of equilibrated samples from earlier checkpoints supplies the negative phase so that only the newest replicas need full simulation.

What would settle it

Train identical RBMs with PTT and PCD on a strongly multimodal discrete dataset (e.g., low-temperature Ising or a clustered genomic set); if PTT test log-likelihood, mode coverage, and two- and three-body correlation errors are not systematically better once both methods are given the same wall-clock budget and thermalization checks, the central claim fails.

Watch

Extended reading notes

Core claim

The authors establish that equilibrium maximum-likelihood training of energy-based models becomes practical once sampling is done by Parallel Trajectory Tempering with reservoir sampling and adaptive optimization: the negative phase stays equilibrated even on highly multimodal or data-scarce scientific data, computational cost remains comparable to Persistent Contrastive Divergence, and the same machinery yields reliable log-likelihoods and thermalization diagnostics. Experiments on Restricted Boltzmann Machines show consistent gains over existing EBM trainers and over state-of-the-art deep generative models on the reported discrete benchmarks.

Load-bearing premise

The learning path stays continuous enough that an adaptive ladder of checkpoints always keeps enough overlap between neighbors for valid exchanges, so earlier frozen samples remain trustworthy sources for the negative phase.

Editorial extensions

If this is right

  • Equilibrium maximum-likelihood training of compact EBMs becomes a practical default rather than a theoretical ideal on discrete scientific data.
  • Log-likelihoods, thermalization times, and equilibrium samples are available during and after training at essentially no extra cost, enabling reliable early stopping and model selection.
  • On multimodal or small tabular datasets, shallow PTT-trained RBMs can outperform larger deep generative models in sample fidelity and overfitting resistance.
  • Failures of persistent chains that look like memorization or mode collapse can be diagnosed directly from replica diffusion across the ladder.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same trajectory-tempering idea may transfer to continuous-energy EBMs if a cheap local sampler replaces Alternating Gibbs Sampling and the adaptive ladder criterion is retuned.
  • Scientific domains that already use pairwise energy models (protein contact prediction, neural population codes) could adopt PTT-trained higher-order EBMs without changing their downstream analysis pipelines.
  • If ladder growth stays O(10) checkpoints on larger problems, PTT could serve as a drop-in negative-phase module inside hybrid generative stacks that still want likelihoods.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript proposes Parallel Trajectory Tempering (PTT) as a practical training algorithm for Energy-Based Models, specialized here to Restricted Boltzmann Machines. Rather than tempering in temperature, PTT builds an adaptive ladder of model checkpoints along the learning trajectory and exchanges configurations between neighbors (Eq. 7), combined with a reservoir of equilibrated samples from frozen earlier checkpoints so that most gradient steps evolve only the last two replicas. Adaptive learning-rate control (cosine similarity of successive gradients plus swap-acceptance safeguards) and recursive partition-function estimation (Eqs. 8–11) are included. On 2D Ising, human genome SNP data, a protein family MSA, neuropixels spike trains, and a medical tabular set, PTT-trained RBMs are reported to outperform PCD-trained RBMs and, on several discrete/tabular benchmarks, Bayesian Flow Networks, with improved mode coverage, higher-order statistics, PRIVET nearest-neighbor diagnostics, contact PPV, and test log-likelihood, while supplying thermalization times and equilibrium samples essentially for free.

Significance. If the claims hold, the work removes a long-standing practical barrier to equilibrium maximum-likelihood training of EBMs on multimodal and data-scarce scientific data, at a cost argued to be comparable to PCD. The Ising finite-size scaling and Binder crossing, PRIVET analysis on HGD, higher-order neural correlations, and protein contact PPV are concrete, falsifiable gains over known failure modes of CD/PCD. Free by-products—τ_exp/τ_int diagnostics, recursive log Z, and post-training equilibrium samples—are genuine methodological strengths for scientific use of EBMs. The contribution is partly engineering (reservoir + adaptive LR on top of prior PTT [30]), but the empirical breadth and the explicit equilibrium-throughout-training goal are valuable for the field.

major comments (3)
  1. [Abstract, §II, Materials & Methods] Abstract and §II assert that reservoir PTT has “computational cost comparable to Persistent Contrastive Divergence” and is a “practical replacement,” on the grounds that checkpoints are added only O(10) times so most updates evolve only the last two replicas. No per-dataset checkpoint counts, fraction of wall-clock or AGS sweeps spent on full-ladder thermalization (20τ_exp) and reservoir fills (N_res=10 N_chains), or matched PCD wall-clock/FLOP tables are reported. Fig. 2c gives ladder size only for Ising; HGD, PF13354, neuropixels, and medical runs have no compute ledger. Without these numbers the cost half of the central claim is unsupported, even if sample quality is strong. Please add a quantitative cost comparison (checkpoints added, total PTT vs PCD sweeps or wall-clock under matched hardware and chain counts) for every main experiment.
  2. [Abstract, §IV Conclusion] All positive results and the continuity/reservoir design are demonstrated for RBMs with Alternating Gibbs Sampling. The abstract and conclusion frame the method as making “equilibrium maximum-likelihood training of EBMs” practical in general. Either restrict the claim to bipartite RBMs (or models with cheap conditional Gibbs) or provide at least one non-RBM EBM experiment (or a clear negative result) showing that the adaptive ladder and reservoir remain valid when local moves mix more slowly. As written, generalization beyond RBMs is an untested load-bearing extrapolation.
  3. [Materials & Methods; Fig. 5] Materials & Methods fix α=0.3, 10 AGS steps per sweep, two swap attempts, N_chains=1000, N_res=10 N_chains, and 20τ_exp thermalization, with LR halving when acceptance drops. Fig. 5 shows PCD is highly sensitive to chain count, but no analogous ablation is given for PTT’s free parameters (α, AGS steps, reservoir size, cosine-similarity rule). A short sensitivity study on at least one multimodal dataset (e.g. HGD or PF13354)—final test LL and τ_exp vs α and AGS steps—would show that equilibrium and quality are not brittle to these choices.
minor comments (6)
  1. [Fig. 3] Fig. 3 caption is corrupted: “Comparison The same behavior is recovered with AIS usingof the sampling quality on HGD.” Clean up and restore the intended caption text.
  2. [Throughout] Several headings and labels show spacing artifacts (“RESUL TS”, “F amily”, “MA TERIALS”), likely from PDF extraction; fix in the camera-ready source.
  3. [Eq. (1), §II] Eq. (1) writes Z θ with an inconsistent space; unify notation for Z_θ / Zθ across Eqs. (1), (8), (10)–(12).
  4. [§III.B–E] State explicitly how BFN and edDCA were trained/hyperparameter-selected (epochs, early stopping, architecture) so that “surpasses SOTA” comparisons are reproducible.
  5. [Materials & Methods; Introduction] Fig. 7 (MNIST data-scarce PCD failure) is valuable but sits only in Materials; a brief forward pointer in the introduction or §III would help readers find the motivating failure mode.
  6. [Fig. 6, Fig. 7] Clarify whether test log-likelihoods for PCD models in Figs. 6–7 use the same PTT/AIS estimator as PTT models, and report uncertainty on LL curves.

Circularity Check

1 steps flagged · score 1.0 of 10

No load-bearing circularity: standard ML objective and classical thermodynamic integration; prior self-citations supply the PTT method, not the empirical claims.

  1. self citation load bearing [§II NEW TRAINING METHOD; Abstract; refs [28],[30]]
    "Motivated by a theoretical analysis of RBM learning dynamics, Parallel Trajectory Tempering (PTT) was recently introduced [30]. Rather than tempering in temperature, PTT exchanges replicas between neighboring models along the learning trajectory (Fig. 1), following Hamiltonian-exchange MCMC [43]. In this work, we make PTT practical for EBM training by combining reservoir sampling with adaptive optimization"

    PTT as a sampling idea is imported from the authors’ own prior work [30] (and related [28]), so the method’s origin is not independent. This is ordinary cumulative research, not circular derivation: the present claims (equilibrium training at PCD-like cost; superior samples/LL vs PCD and BFN on Ising, HGD, PF13354, neuropixels, medical data) are empirical and checked against external observables, not forced by the self-citation.

full rationale

The paper’s training objective is ordinary maximum-likelihood (Eqs. 2–4); the negative-phase estimator is MCMC, and the partition-function recursion (Eqs. 8/10–11) is classical thermodynamic integration / free-energy perturbation along an adaptive ladder, not a tautology that forces the reported likelihoods or sample-quality metrics. Model selection uses test LL, but generative quality is judged against external, non-fitted benchmarks (exact 2D Ising scaling and Binder crossing; PRIVET NN-distance reference; edDCA contact PPV; BFN samples; empirical two-/three-body correlations and P(K)). Self-citations ([28],[30],[7],[3]) introduce PTT, hierarchical free-energy clustering, and coupling extraction—method and analysis tools from overlapping authors—but they are not used as uniqueness theorems that forbid alternatives or as fitted inputs renamed as predictions. Cost-comparability to PCD is an algorithmic O(10)-checkpoint argument rather than a measured ledger; that is an evidence gap, not circularity. No step reduces a claimed prediction to its own definition or fit by construction. Score 1 reflects only non-load-bearing method self-citation.

Assumptions & free parameters 5 free parameters · 5 assumptions · 1 invented entities

The central claim rests on standard MCMC/ML machinery plus engineering choices that keep a trajectory ladder equilibrated. Load-bearing modeling choices are RBM+AGS expressivity and the smoothness of the optimization path; numerical knobs (α, reservoir size, τ multipliers, LR schedule) are hand-set but monitored. No new physical entities are postulated.

free parameters (5)
  • swap_acceptance_threshold_alpha = 0.3 (fixed)
    New checkpoints are frozen when swap acceptance with the previous replica falls below α; also used to reject steps and halt if any ladder edge drops below α−0.2. Directly controls ladder density and claimed equilibrium.
  • initial_and_adaptive_learning_rate = init 1e-3; adaptive
    Initialized at 1e-3, halved on failed swaps, modulated by cosine similarity of successive gradients (per RBM parameter block). Affects whether the trajectory stays continuous enough for PTT.
  • AGS_steps_per_PTT_sweep_and_swap_attempts = k=10; 2 swap attempts
    Each sweep uses 10 alternating Gibbs updates and two swap attempts between reservoir and last checkpoints; sets mixing per gradient step and compute parity vs PCD.
  • thermalization_multipliers_and_reservoir_size = 20 τ_exp; 2 τ_int; N_res=10×1000
    Burn-in ≥20 τ_exp, store every 2 τ_int, N_res=10 N_chains with N_chains=1000. Defines what counts as equilibrium samples for training and evaluation.
  • RBM_hidden_unit_counts_and_early_stopping = dataset-dependent; max test LL checkpoint
    Architecture size (e.g. N_h=2L for Ising) and model selection at max test LL are chosen per experiment and affect capacity and reported generalization.
assumptions (5)
  • standard math Metropolis exchanges on energy differences between neighboring trajectory checkpoints yield detailed balance on the product extended ensemble.
    Eq. 7 and Hamiltonian-exchange MCMC background; standard if proposals are symmetric and models are fixed during a sweep.
  • domain assumption Unbiased ML gradients require equilibrium model samples; out-of-equilibrium persistent chains bias learning on multimodal or scarce data.
    §I.A and introduction; motivates replacing CD/PCD negative phase with PTT.
  • domain assumption RBM training trajectories are smooth enough that an adaptive checkpoint ladder maintains neighbor overlap without needing a full temperature ladder.
    §II and motivation from prior spectral/phase-transition analyses; core reason PTT can be cheaper than PT.
  • domain assumption Replica index autocorrelation on the ladder diagnoses global thermalization (uniform visitation, τ_exp / τ_int).
    Materials & Methods citing spin-glass PT diagnostics [47,68]; used as stopping and reservoir protocol.
  • standard math Recursive reweighting Z_t = ⟨e^{E_{t-1}-E_t}⟩_{t-1} Z_{t-1} gives accurate log-likelihood when neighbors overlap strongly.
    Eqs. 8–12; low-variance only because α enforces overlap—tied to free parameter α.
invented entities (1)
  • Parallel Trajectory Tempering (PTT) training loop with equilibrium reservoir independent evidence
    purpose: Replace temperature tempering during learning by exchanging along frozen model checkpoints and only simulating the newest replicas plus a stored reservoir.
    PTT concept credited to prior author work [30]; this paper's entity is the practical reservoir+adaptive-LR training algorithm claimed to match PCD cost.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering." pith.science (2026). https://pith.science/paper/2UW6AF4V

@misc{pith2026260727077,
  author       = {Pith},
  title        = {Pith review of: Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2UW6AF4V}},
  note         = {Machine review of arXiv:2607.27077}
}
read the original abstract

Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the continuity of the optimization path to maintain equilibrium sampling throughout learning. This enables stable and fast training on highly multimodal and data-scarce scientific datasets. Combined with reservoir sampling and adaptive optimization, PTT has a computational cost comparable to Persistent Contrastive Divergence, making it a practical replacement for standard training methods. It also provides direct estimates of thermalization times, equilibrium samples from trained models, and accurate log-likelihoods at essentially no additional cost. Experiments on Restricted Boltzmann Machines show that PTT consistently outperforms existing EBM training approaches. On discrete tabular data, it also surpasses state-of-the-art deep generative models, yielding higher-quality samples and greater robustness to overfitting and limited data. Our results make equilibrium maximum-likelihood training of EBMs practical and computationally efficient.

Figures

Figures reproduced from arXiv: 2607.27077 by the authors.

Figure 1
Figure 1. FIG. 1. Schematic illustration of the Parallel Trajectory Tem [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 6
Figure 6. Figure 6: FIG. 6 [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 7
Figure 7. Figure 7: FIG. 7 [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

69 extracted references · 5 linked inside Pith

  1. [30]

    B´ ereux, A

    N. B´ ereux, A. Decelle, C. Furtlehner, L. Rosset, and B. Seoane, Fast training and sampling of restricted Boltz- mann machines, in13th International Conference on Learning Representations-ICLR 2025(2025)

  2. [1]

    Hinton, Nobel lecture: Boltzmann machines, Reviews of Modern Physics97, 030502 (2025)

    G. Hinton, Nobel lecture: Boltzmann machines, Reviews of Modern Physics97, 030502 (2025)

  3. [2]

    Bulso and Y

    N. Bulso and Y. Roudi, Restricted Boltzmann machines as models of interacting variables, Neural Computation 33, 2646 (2021)

  4. [3]

    Decelle, A

    A. Decelle, A. d. J. Navas G´ omez, and B. Seoane, Infer- ring higher-order couplings with neural networks, Physi- cal Review Letters135, 207301 (2025)

  5. [4]

    Decelle, A

    A. Decelle, A. d. J. N. G´ omez, and B. Seoane, Distri- butional simplicity bias and effective convexity in energy based models, arXiv preprint arXiv:2605.07844 (2026)

  6. [5]

    Tubiana and R

    J. Tubiana and R. Monasson, Emergence of composi- tional representations in restricted Boltzmann machines, Physical review letters118, 138301 (2017)

  7. [6]

    Tubiana, S

    J. Tubiana, S. Cocco, and R. Monasson, Learning protein constitutive motifs from sequence data, Elife8, e39397 (2019)

  8. [7]

    Decelle, B

    A. Decelle, B. Seoane, and L. Rosset, Unsupervised hi- erarchical clustering using the learning dynamics of re- stricted Boltzmann machines, Phys. Rev. E108, 014110 (2023)

Show all 69 references
  1. [8]

    di Sarra, B

    G. di Sarra, B. Bravi, and Y. Roudi, The unbearable lightness of restricted Boltzmann machines: Theoretical insights and biological applications, Europhysics Letters 149, 21002 (2025)

  2. [9]

    Bravi, J

    B. Bravi, J. Tubiana, S. Cocco, R. Monasson, T. Mora, and A. M. Walczak, Rbm-mhc: a semi-supervised machine-learning method for sample-specific prediction of antigen presentation by hla-i alleles, Cell systems12, 195 (2021)

  3. [10]

    Bravi, Development and use of machine learning al- gorithms in vaccine target selection, npj Vaccines9, 15 (2024)

    B. Bravi, Development and use of machine learning al- gorithms in vaccine target selection, npj Vaccines9, 15 (2024)

  4. [11]

    T. L. van der Plas, J. Tubiana, G. Le Goc, G. Mi- gault, M. Kunst, H. Baier, V. Bormuth, B. Englitz, and G. Debr´ egeas, Neural assemblies uncovered by genera- tive modeling explain whole-brain activity statistics and reflect structural connectivity, Elife12, e83139 (2023)

  5. [12]

    B´ ereux, G

    N. B´ ereux, G. Catania, A. Decelle, F. Mignacco, A. d. J. N. G´ omez, and B. Seoane, Uncovering statistical struc- ture in large-scale neural activity with restricted Boltz- mann machines, arXiv preprint arXiv:2603.11032 (2026)

  6. [13]

    Dommanget-Kott, J

    M. Dommanget-Kott, J. Fernandez-de Cossio-Diaz, G. Faye-B´ edrin, G. Debr´ egeas, and V. Bormuth, Cross- individual translation of spontaneous zebrafish brain activity through a shared latent representation, Pro- ceedings of the National Academy of Sciences123, e2529064123 (2026)

  7. [14]

    R. G. Melko, G. Carleo, J. Carrasquilla, and J. I. Cirac, Restricted Boltzmann machines in quantum physics, Na- ture Physics15, 887 (2019)

  8. [15]

    Barra, G

    A. Barra, G. Genovese, P. Sollich, and D. Tantari, Phase diagram of restricted Boltzmann machines and general- ized hopfield networks with arbitrary priors, Physical Re- view E97, 022310 (2018)

  9. [16]

    Decelle and C

    A. Decelle and C. Furtlehner, Restricted Boltzmann ma- chine: Recent advances and mean-field theory, Chinese Physics B30, 040202 (2021)

  10. [17]

    Carleo and M

    G. Carleo and M. Troyer, Solving the quantum many- body problem with artificial neural networks, Science 355, 602 (2017)

  11. [18]

    Nomura, A

    Y. Nomura, A. S. Darmawan, Y. Yamaji, and M. Imada, Restricted Boltzmann machine learning for solving strongly correlated quantum systems, Physical Review B96, 205152 (2017)

  12. [19]

    Decelle, C

    A. Decelle, C. Furtlehner, and B. Seoane, Equilibrium and non-equilibrium regimes in the learning of restricted Boltzmann machines, Advances in Neural Information Processing Systems34, 5345 (2021)

  13. [20]

    R. Liao, S. Kornblith, M. Ren, D. J. Fleet, and G. Hinton, Gaussian-Bernoulli rbms without tears, arXiv preprint arXiv:2210.10318 (2022)

  14. [21]

    Decelle, G

    A. Decelle, G. Fissore, and C. Furtlehner, Spectral dy- namics of learning in restricted Boltzmann machines, Eu- rophysics Letters119, 60001 (2017)

  15. [22]

    Decelle, G

    A. Decelle, G. Fissore, and C. Furtlehner, Thermody- namics of restricted Boltzmann machines and related learning dynamics, Journal of Statistical Physics172, 1576 (2018)

  16. [23]

    Bachtis, G

    D. Bachtis, G. Biroli, A. Decelle, and B. Seoane, Cas- cade of phase transitions in the training of energy-based models, NeurIPS (2024), arXiv:2405.14689 (2024)

  17. [24]

    Krause, A

    O. Krause, A. Fischer, and C. Igel, Algorithms for es- timating the partition function of restricted Boltzmann machines, Artificial Intelligence278, 103195 (2020)

  18. [25]

    Nijkamp, M

    E. Nijkamp, M. Hill, S.-C. Zhu, and Y. N. Wu, Learn- ing non-convergent non-persistent short-run mcmc to- ward energy-based model, Advances in Neural Informa- tion Processing Systems32(2019)

  19. [26]

    Nijkamp, M

    E. Nijkamp, M. Hill, T. Han, S.-C. Zhu, and Y. N. Wu, On the anatomy of MCMC-based maximum like- lihood learning of energy-based models, inProceedings of 8 the AAAI Conference on Artificial Intelligence, Vol. 34 (2020) pp. 5272–5280

  20. [27]

    Agoritsas, G

    E. Agoritsas, G. Catania, A. Decelle, and B. Seoane, Explaining the effects of non-convergent sampling in the training of energy-based models, arXiv preprint arXiv:2301.09428 (2023)

  21. [28]

    B´ ereux, A

    N. B´ ereux, A. Decelle, C. Furtlehner, and B. Seoane, Learning a restricted Boltzmann machine using biased Monte Carlo sampling, SciPost Physics14, 032 (2023)

  22. [29]

    Carbone, A

    A. Carbone, A. Decelle, L. Rosset, and B. Seoane, Fast and functional structured data generators rooted in out- of-equilibrium physics, IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  23. [31]

    G. E. Hinton, Training products of experts by minimiz- ing contrastive divergence, Neural computation14, 1771 (2002)

  24. [32]

    T. Tieleman, Training restricted Boltzmann machines us- ing approximations to the likelihood gradient, inPro- ceedings of the 25th international conference on Machine learning(2008) pp. 1064–1071

  25. [33]

    Decelle and C

    A. Decelle and C. Furtlehner, Exact training of restricted Boltzmann machines on intrinsically low dimensional data, Physical Review Letters127, 158303 (2021)

  26. [34]

    Krause, A

    O. Krause, A. Fischer, and C. Igel, Population- contrastive-divergence: Does consistency help with RBM training?, Pattern Recognition Letters102, 1 (2018)

  27. [35]

    Carbone, M

    D. Carbone, M. Hua, S. Coste, and E. Vanden-Eijnden, Efficient training of energy-based models using Jarzyn- ski equality, Advances in Neural Information Processing Systems36(2024)

  28. [36]

    Marinari and G

    E. Marinari and G. Parisi, Simulated tempering: a new Monte Carlo scheme, Europhysics letters19, 451 (1992)

  29. [37]

    Lyubartsev, A

    A. Lyubartsev, A. Martsinovski, S. Shevkunov, and P. Vorontsov-Velyaminov, New approach to Monte Carlo calculation of the free energy: Method of expanded en- sembles, The Journal of chemical physics96, 1776 (1992)

  30. [38]

    C. J. Geyer and E. A. Thompson, Annealing Markov chain Monte Carlo with applications to ancestral infer- ence, Journal of the American Statistical Association90, 909 (1995)

  31. [39]

    M. Tesi, E. Janse van Rensburg, E. Orlandini, and S. Whittington, Monte Carlo study of the interacting self-avoiding walk model in three dimensions, Journal of statistical physics82, 155 (1996)

  32. [40]

    Hukushima and K

    K. Hukushima and K. Nemoto, Exchange Monte Carlo method and application to spin glass simulations, Journal of the Physical Society of Japan65, 1604 (1996)

  33. [41]

    R. R. Salakhutdinov, Learning in Markov random fields using tempered transitions, Advances in neural informa- tion processing systems22(2009)

  34. [42]

    Desjardins, A

    G. Desjardins, A. Courville, Y. Bengio, P. Vincent, and O. Delalleau, Tempered markov chain Monte Carlo for training of restricted Boltzmann machines, inProceed- ings of the thirteenth international conference on arti- ficial intelligence and statistics(JMLR Workshop and Confe...

  35. [43]

    Rosta, M

    E. Rosta, M. Nowotny, W. Yang, and G. Hummer, Cat- alytic mechanism of rna backbone cleavage by ribonu- clease h from quantum mechanics/molecular mechanics simulations, Journal of the American Chemical Society 133, 8934 (2011)

  36. [44]

    Smolensky, In parallel distributed processing: Volume 1 by d

    P. Smolensky, In parallel distributed processing: Volume 1 by d. rumelhart and j. mclelland (MIT Press, 1986) Chap. 6: Information Processing in Dynamical Systems: Foundations of Harmony Theory

  37. [45]

    Le Roux and Y

    N. Le Roux and Y. Bengio, Representational power of restricted Boltzmann machines and deep belief networks, Neural computation20, 1631 (2008)

  38. [46]

    Decelle, C

    A. Decelle, C. Furtlehner, A. d. J. Navas G´ omez, and B. Seoane, Inferring effective couplings with restricted Boltzmann machines, SciPost Physics16, 095 (2024)

  39. [47]

    Alvarez Ba˜ nos, A

    R. Alvarez Ba˜ nos, A. Cruz, L. Fernandez, J. Gil-Narvion, A. Gordillo-Guerrero, M. Guidetti, A. Maiorano, F. Man- tovani, E. Marinari, V. Martin-Mayor,et al., Nature of the spin-glass phase at experimental length scales, Jour- nal of Statistical Mechanics: Theory and Experime...

  40. [48]

    Graves, R

    A. Graves, R. K. Srivastava, T. Atkinson, and F. Gomez, Bayesian flow networks, arXiv preprint arXiv:2308.07037 (2023)

  41. [49]

    Yevick and R

    D. Yevick and R. Melko, The accuracy of restricted Boltzmann machine models of ising systems, Computer Physics Communications258, 107518 (2021)

  42. [50]

    Gu and K

    J. Gu and K. Zhang, Thermodynamics of the ising model encoded in restricted Boltzmann machines, Entropy24, 1701 (2022)

  43. [51]

    M. A. Valle, The capabilities of Boltzmann machines to detect and reconstruct ising system’s configurations from a given temperature, Entropy25, 1649 (2023)

  44. [52]

    R. H. Swendsen and J.-S. Wang, Nonuniversal critical dynamics in Monte Carlo simulations, Physical review letters58, 86 (1987)

  45. [53]

    D. J. Amit and V. Martin-Mayor,Field theory, the renor- malization group, and critical phenomena: graphs to computers(World Scientific Publishing Company, 2005)

  46. [54]

    Yelmen, A

    B. Yelmen, A. Decelle, L. Ongaro, D. Marnetto, C. Tal- lec, F. Montinaro, C. Furtlehner, L. Pagani, and F. Jay, Creating artificial human genomes using generative neu- ral networks, PLoS genetics17, e1009303 (2021)

  47. [55]

    G. P. Consortiumet al., A global reference for human genetic variation, Nature526, 68 (2015)

  48. [56]

    Yelmen, A

    B. Yelmen, A. Decelle, L. L. Boulos, A. Szatkownik, C. Furtlehner, G. Charpiat, and F. Jay, Deep convo- lutional and conditional neural networks for large-scale genomic data generation, PLOS Computational Biology 19, e1011584 (2023)

  49. [57]

    Yelmen and F

    B. Yelmen and F. Jay, An overview of deep generative models in functional and evolutionary genomics, Annual Review of Biomedical Data Science6, 173 (2023)

  50. [58]

    Szatkownik, A

    A. Szatkownik, A. Decelle, B. Seoane, N. B´ ereux, L. Planche, G. Charpiat, B. Yelmen, F. Jay, and C. Furtlehner, PRIVET: Privacy metric based on ex- treme value theory, arXiv preprint arXiv:2510.24233 (2025)

  51. [59]

    Szatkownik, L

    A. Szatkownik, L. Planche, M. Demeulle, T. Chambe, M. C. ´Avila-Arcos, E. Huerta-Sanchez, C. Furtlehner, G. Charpiat, F. Jay, and B. Yelmen, Diffusion-based ar- tificial genomes and their usefulness for local ancestry inference, bioRxiv , 2024 (2024)

  52. [60]

    A. P. Muntoni, A. Pagnani, M. Weigt, and F. Zamponi, adabmDCA: adaptive Boltzmann machine learning for biological sequences, BMC bioinformatics22, 528 (2021)

  53. [61]

    Barrat-Charlaix, A

    P. Barrat-Charlaix, A. P. Muntoni, K. Shimagaki, 9 M. Weigt, and F. Zamponi, Sparse generative modeling via parameter reduction of Boltzmann machines: appli- cation to protein-sequence families, Physical Review E 104, 024407 (2021)

  54. [62]

    Rosset, R

    L. Rosset, R. Netti, A. P. Muntoni, M. Weigt, and F. Zamponi, adabmdca 2.0—a flexible but easy-to-use package for direct coupling analysis, inProtein Evolution: Methods and Protocols(Springer, 2026) pp. 83–104

  55. [63]

    Rubner, C

    Y. Rubner, C. Tomasi, and L. J. Guibas, A metric for dis- tributions with applications to image databases, inSixth international conference on computer vision (IEEE Cat. No. 98CH36271)(IEEE, 1998) pp. 59–66

  56. [64]

    N. A. Steinmetz, C. Aydin, A. Lebedeva, M. Okun, M. Pachitariu, M. Bauza, M. Beau, J. Bhagat, C. B¨ ohm, M. Broux, S. Chen, J. Colonell, R. J. Gardner, B. Karsh, F. Kloosterman, D. Kostadinov, C. Mora-Lopez, J. O’Callaghan, J. Park, J. Putzeys, B. Sauerbrei, R. J. J. van Daal,...

  57. [65]

    Medical recommendation system,https: //www.kaggle.com/datasets/joymarhew/ medical-reccomadation-dataset/data(2023), ac- cessed 2026-07-24

  58. [66]

    L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veera- machaneni, Modeling tabular data using conditional gan, Advances in neural information processing systems32 (2019)

  59. [67]

    Kotelnikov, D

    A. Kotelnikov, D. Baranchuk, I. Rubachev, and A. Babenko, Tabddpm: Modelling tabular data with dif- fusion models, inInternational conference on machine learning(PMLR, 2023) pp. 17564–17579

  60. [68]

    R. A. Banos, A. Cruz, L. Fernandez, J. Gil-Narvion, A. Gordillo-Guerrero, M. Guidetti, A. Maiorano, F. Man- tovani, E. Marinari, V. Martin-Mayor,et al., Nature of the spin-glass phase at experimental length scales, Jour- nal of Statistical Mechanics: Theory and Experiment 2010...

  61. [69]

    LeCun, L

    Y. LeCun, L. Bottou, Y. Bengio, P. Haffner,et al., Gradient-based learning applied to document recogni- tion, Proceedings of the IEEE86, 2278 (1998). MA TERIALS AND METHODS Hyperparameters All models are trained using two PTT sweep attempts between the reservoir and the last t...

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.