Pith. sign in

REVIEW 1 cited by

Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.11867 v1 pith:TGCAGLLO submitted 2020-04-24 cs.CL

classification cs.CL
keywords translationmultilingualzero-shotlanguagemodelsperformancebilingualmachine
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Massively multilingual models for neural machine translation (NMT) are theoretically attractive, but often underperform bilingual models and deliver poor zero-shot translations. In this paper, we explore ways to improve them. We argue that multilingual NMT requires stronger modeling capacity to support language pairs with varying typological characteristics, and overcome this bottleneck via language-specific components and deepening NMT architectures. We identify the off-target translation issue (i.e. translating into a wrong target language) as the major source of the inferior zero-shot performance, and propose random online backtranslation to enforce the translation of unseen training language pairs. Experiments on OPUS-100 (a novel multilingual dataset with 100 languages) show that our approach substantially narrows the performance gap with bilingual models in both one-to-many and many-to-many settings, and improves zero-shot performance by ~10 BLEU, approaching conventional pivot-based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gradients: When Markets Meet Fine-tuning -- A Distributed Approach to Model Optimisation

    cs.AI 2025-06 reject novelty 6.0 of 10

    Gradients reports that competitive, reward-driven fine-tuning beats centralized AutoML in 82 to 100 percent of comparisons, but its evaluation does not isolate competition from a much larger compute budget.

Pith tools