Pith. sign in

REVIEW 1 cited by

Dual Training of Energy-Based Models with Overparametrized Shallow Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.05134 v2 pith:YWGBSHP3 submitted 2021-07-11 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords dualenergymodelsparticlestrainingalgorithmebmsenergy-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Energy-based models (EBMs) are generative models that are usually trained via maximum likelihood estimation. This approach becomes challenging in generic situations where the trained energy is non-convex, due to the need to sample the Gibbs distribution associated with this energy. Using general Fenchel duality results, we derive variational principles dual to maximum likelihood EBMs with shallow overparametrized neural network energies, both in the feature-learning and lazy linearized regimes. In the feature-learning regime, this dual formulation justifies using a two time-scale gradient ascent-descent (GDA) training algorithm in which one updates concurrently the particles in the sample space and the neurons in the parameter space of the energy. We also consider a variant of this algorithm in which the particles are sometimes restarted at random samples drawn from the data set, and show that performing these restarts at every iteration step corresponds to score matching training. These results are illustrated in simple numerical experiments, which indicates that GDA performs best when features and particles are updated using similar time scales.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Jarzynski reweighting, which estimates normalization constants from out-of-equilibrium paths, is shown to apply to a broad class of sampling kernels, including drift-based stochastic interpolants and RBM Gibbs samplin...

Pith tools