REVIEW 2 cited by
Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While recent work has shown that scores from models trained by the ubiquitous masked language modeling (MLM) objective effectively discriminate probable from improbable sequences, it is still an open question if these MLMs specify a principled probability distribution over the space of possible sequences. In this paper, we interpret MLMs as energy-based sequence models and propose two energy parametrizations derivable from the trained MLMs. In order to draw samples correctly from these models, we develop a tractable sampling scheme based on the Metropolis--Hastings Monte Carlo algorithm. In our approach, samples are proposed from the same masked conditionals used for training the masked language models, and they are accepted or rejected based on their energy values according to the target distribution. We validate the effectiveness of the proposed parametrizations by exploring the quality of samples drawn from these energy-based models for both open-ended unconditional generation and a conditional generation task of machine translation. We theoretically and empirically justify our sampling algorithm by showing that the masked conditionals on their own do not yield a Markov chain whose stationary distribution is that of our target distribution, and our approach generates higher quality samples than other recently proposed undirected generation approaches (Wang et al., 2019, Ghazvininejad et al., 2019).
Forward citations
Cited by 2 Pith papers
-
GenPlan: Generative Sequence Models as Adaptive Planners
A discrete-flow sequence model with energy and entropy guidance enables adaptive planning that beats LEAP and Decision Transformer in BabyAI adaptive benchmarks.
-
Adaptformer: Sequence models as adaptive iterative planners
Adaptformer extends LEAP-style energy-based planning with a learned intrinsic sub-goal curriculum and entropy-regularized stochastic policy, enabling generalization to multi-goal and other out-of-distribution missions.
Discussion (0). Continue with ORCID to comment.