Pith. sign in

REVIEW 2 cited by

Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.02736 v2 pith:4P5P4J65 submitted 2021-06-04 cs.LG cs.CL

classification cs.LGcs.CL
keywords modelsmaskeddistributionsamplesenergygenerationlanguagemlms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While recent work has shown that scores from models trained by the ubiquitous masked language modeling (MLM) objective effectively discriminate probable from improbable sequences, it is still an open question if these MLMs specify a principled probability distribution over the space of possible sequences. In this paper, we interpret MLMs as energy-based sequence models and propose two energy parametrizations derivable from the trained MLMs. In order to draw samples correctly from these models, we develop a tractable sampling scheme based on the Metropolis--Hastings Monte Carlo algorithm. In our approach, samples are proposed from the same masked conditionals used for training the masked language models, and they are accepted or rejected based on their energy values according to the target distribution. We validate the effectiveness of the proposed parametrizations by exploring the quality of samples drawn from these energy-based models for both open-ended unconditional generation and a conditional generation task of machine translation. We theoretically and empirically justify our sampling algorithm by showing that the masked conditionals on their own do not yield a Markov chain whose stationary distribution is that of our target distribution, and our approach generates higher quality samples than other recently proposed undirected generation approaches (Wang et al., 2019, Ghazvininejad et al., 2019).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenPlan: Generative Sequence Models as Adaptive Planners

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A discrete-flow sequence model with energy and entropy guidance enables adaptive planning that beats LEAP and Decision Transformer in BabyAI adaptive benchmarks.

  2. Adaptformer: Sequence models as adaptive iterative planners

    cs.RO 2024-11 conditional novelty 6.0 of 10

    Adaptformer extends LEAP-style energy-based planning with a learned intrinsic sub-goal curriculum and entropy-regularized stochastic policy, enabling generalization to multi-goal and other out-of-distribution missions.

Pith tools