Pith. sign in

REVIEW 3 cited by

Investigating Backtranslation in Neural Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.06189 v1 pith:25NKHBGV submitted 2018-04-17 cs.CL

classification cs.CL
keywords dataparalleltranslationback-translatedsystemsavailablebacktranslationbecome
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A prerequisite for training corpus-based machine translation (MT) systems -- either Statistical MT (SMT) or Neural MT (NMT) -- is the availability of high-quality parallel data. This is arguably more important today than ever before, as NMT has been shown in many studies to outperform SMT, but mostly when large parallel corpora are available; in cases where data is limited, SMT can still outperform NMT. Recently researchers have shown that back-translating monolingual data can be used to create synthetic parallel corpora, which in turn can be used in combination with authentic parallel data to train a high-quality NMT system. Given that large collections of new parallel text become available only quite rarely, backtranslation has become the norm when building state-of-the-art NMT systems, especially in resource-poor scenarios. However, we assert that there are many unknown factors regarding the actual effects of back-translated data on the translation capabilities of an NMT model. Accordingly, in this work we investigate how using back-translated data as a training corpus -- both as a separate standalone dataset as well as combined with human-generated parallel data -- affects the performance of an NMT model. We use incrementally larger amounts of back-translated data to train a range of NMT systems for German-to-English, and analyse the resulting translation performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Back-Translation with Uncertainty-based Confidence Estimation

    cs.CL 2019-08 accept novelty 6.0 of 10

    Uncertainty-based confidence estimation, computed with Monte Carlo Dropout, improves back-translation for NMT by weighting synthetic sentence pairs and reweighting attention, yielding consistent BLEU gains on Chinese-...

  2. On The Evaluation of Machine Translation Systems Trained With Back-Translation

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Back-translation produces more fluent, human-preferred output even when BLEU is flat, so evaluation should combine BLEU with a language model score.

  3. Learning Credible Deep Neural Networks with Rationale Regularization

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Regularizing a deep network's local explanations to match human rationales during training improves accuracy on out-of-distribution text data without hurting in-distribution test accuracy.

Pith tools