Pith. sign in

REVIEW 1 cited by

Fully Synthetic Data Improves Neural Machine Translation with Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15455 v3 pith:XCZZUATK submitted 2020-12-31 cs.CL

classification cs.CL
keywords languagemonolingualsourcetargetdatatesttranslationcorpus
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper explores augmenting monolingual data for knowledge distillation in neural machine translation. Source language monolingual text can be incorporated as a forward translation. Interestingly, we find the best way to incorporate target language monolingual text is to translate it to the source language and round-trip translate it back to the target language, resulting in a fully synthetic corpus. We find that combining monolingual data from both source and target languages yields better performance than a corpus twice as large only in one language. Moreover, experiments reveal that the improvement depends upon the provenance of the test set. If the test set was originally in the source language (with the target side written by translators), then forward translating source monolingual data matters. If the test set was originally in the target language (with the source written by translators), then incorporating target monolingual data matters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Back translation with dataset selection raises English-Luganda NMT BLEU scores by about 10 points on a new test set, though the comparison with previous benchmarks is not on a shared test set.

Pith tools