Pith. sign in

REVIEW 1 cited by

Efficient Training of Energy-Based Models Using Jarzynski Equality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.19414 v2 pith:KGZAH4FS submitted 2023-05-30 cs.LG cond-mat.dis-nncs.NAmath.NAmath.PR

classification cs.LGcond-mat.dis-nncs.NAmath.NAmath.PR
keywords algorithmdistributionmodelmodelssamplingcomputationcontrastivecross-entropy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Energy-based models (EBMs) are generative models inspired by statistical physics with a wide range of applications in unsupervised learning. Their performance is best measured by the cross-entropy (CE) of the model distribution relative to the data distribution. Using the CE as the objective for training is however challenging because the computation of its gradient with respect to the model parameters requires sampling the model distribution. Here we show how results for nonequilibrium thermodynamics based on Jarzynski equality together with tools from sequential Monte-Carlo sampling can be used to perform this computation efficiently and avoid the uncontrolled approximations made using the standard contrastive divergence algorithm. Specifically, we introduce a modification of the unadjusted Langevin algorithm (ULA) in which each walker acquires a weight that enables the estimation of the gradient of the cross-entropy at any step during GD, thereby bypassing sampling biases induced by slow mixing of ULA. We illustrate these results with numerical experiments on Gaussian mixture distributions as well as the MNIST dataset. We show that the proposed approach outperforms methods based on the contrastive divergence algorithm in all the considered situations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    An abstract-only MoE paper claiming 97.5% activation prediction accuracy and a 17% to 72% cache hit rate gain, whose full text is a different paper on functional equations, making the results unverifiable.

Pith tools