Pith. sign in

REVIEW

Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.11861 v3 pith:V3ES232P submitted 2019-09-26 cs.CL

classification cs.CL
keywords datasetlarge-scalemachineneuralpairsperformancepretrainingsentence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we investigate the problem of training neural machine translation (NMT) systems with a dataset of more than 40 billion bilingual sentence pairs, which is larger than the largest dataset to date by orders of magnitude. Unprecedented challenges emerge in this situation compared to previous NMT work, including severe noise in the data and prohibitively long training time. We propose practical solutions to handle these issues and demonstrate that large-scale pretraining significantly improves NMT performance. We are able to push the BLEU score of WMT17 Chinese-English dataset to 32.3, with a significant performance boost of +3.2 over existing state-of-the-art results.

Discussion (0). Continue with ORCID to comment.

Pith tools