Pith. sign in

REVIEW 1 cited by

The Impact of Syntactic and Semantic Proximity on Machine Translation with Back-Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.18031 v1 pith:CLEYRIRB submitted 2024-03-26 cs.CL

classification cs.CL
keywords languagesback-translationunsupervisedsemanticacrossmachinemethodsuccess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Unsupervised on-the-fly back-translation, in conjunction with multilingual pretraining, is the dominant method for unsupervised neural machine translation. Theoretically, however, the method should not work in general. We therefore conduct controlled experiments with artificial languages to determine what properties of languages make back-translation an effective training method, covering lexical, syntactic, and semantic properties. We find, contrary to popular belief, that (i) parallel word frequency distributions, (ii) partially shared vocabulary, and (iii) similar syntactic structure across languages are not sufficient to explain the success of back-translation. We show however that even crude semantic signal (similar lexical fields across languages) does improve alignment of two languages through back-translation. We conjecture that rich semantic dependencies, parallel across languages, are at the root of the success of unsupervised methods based on back-translation. Overall, the success of unsupervised machine translation was far from being analytically guaranteed. Instead, it is another proof that languages of the world share deep similarities, and we hope to show how to identify which of these similarities can serve the development of unsupervised, cross-linguistic tools.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Code-switching hurts LLM comprehension when non-English tokens enter English text, but inserting English into other languages often improves accuracy; fine-tuning mitigates losses more reliably than prompting.

Pith tools