Jarzynski reweighting, which estimates normalization constants from out-of-equilibrium paths, is shown to apply to a broad class of sampling kernels, including drift-based stochastic interpolants and RBM Gibbs sampling, with weights that vanish in the continuous-time limit.
Dual Training of Energy-Based Models with Overparametrized Shallow Neural Networks
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Energy-based models (EBMs) are generative models that are usually trained via maximum likelihood estimation. This approach becomes challenging in generic situations where the trained energy is non-convex, due to the need to sample the Gibbs distribution associated with this energy. Using general Fenchel duality results, we derive variational principles dual to maximum likelihood EBMs with shallow overparametrized neural network energies, both in the feature-learning and lazy linearized regimes. In the feature-learning regime, this dual formulation justifies using a two time-scale gradient ascent-descent (GDA) training algorithm in which one updates concurrently the particles in the sample space and the neurons in the parameter space of the energy. We also consider a variant of this algorithm in which the particles are sometimes restarted at random samples drawn from the data set, and show that performing these restarts at every iteration step corresponds to score matching training. These results are illustrated in simple numerical experiments, which indicates that GDA performs best when features and particles are updated using similar time scales.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels
Jarzynski reweighting, which estimates normalization constants from out-of-equilibrium paths, is shown to apply to a broad class of sampling kernels, including drift-based stochastic interpolants and RBM Gibbs sampling, with weights that vanish in the continuous-time limit.