Pith. sign in

REVIEW 2 cited by

Mixture Models for Diverse Machine Translation: Tricks of the Trade

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.07816 v2 pith:U4EABHPV submitted 2019-02-20 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelsmixturediverselatentmachinetranslationvariablediversity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Mixture models trained via EM are among the simplest, most widely used and well understood latent variable models in the machine learning literature. Surprisingly, these models have been hardly explored in text generation applications such as machine translation. In principle, they provide a latent variable to control generation and produce a diverse set of hypotheses. In practice, however, mixture models are prone to degeneracies---often only one component gets trained or the latent variable is simply ignored. We find that disabling dropout noise in responsibility computation is critical to successful training. In addition, the design choices of parameterization, prior distribution, hard versus soft EM and online versus offline assignment can dramatically affect model performance. We develop an evaluation protocol to assess both quality and diversity of generations against multiple references, and provide an extensive empirical study of several mixture model variants. Our analysis shows that certain types of mixture models are more robust and offer the best trade-off between translation quality and diversity compared to variational models and diverse decoding approaches.\footnote{Code to reproduce the results in this paper is available at \url{https://github.com/pytorch/fairseq}}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mixture Content Selection for Diverse Sequence Generation

    cs.CL 2019-09 conditional novelty 6.0 of 10

    A mixture-of-experts content selector that masks different input tokens for each generated sequence improves diversity and accuracy in question generation and summarization.

  2. Optimize TSK Fuzzy Systems for Classification Problems: Mini-Batch Gradient Descent with Uniform Regularization and Batch Normalization

    cs.LG 2019-08 conditional novelty 5.0 of 10

    Adding batch normalization and a uniform-firing regularizer to mini-batch training improves TSK fuzzy classification accuracy on 12 UCI datasets, though the combined gain over the regularizer alone is not statisticall...

Pith tools