pith. sign in

arxiv: 1808.05700 · v1 · pith:TZLBPWQOnew · submitted 2018-08-16 · 💻 cs.CL

Augmenting Statistical Machine Translation with Subword Translation of Out-of-Vocabulary Words

classification 💻 cs.CL
keywords translationmachinestatisticalwordsevaluatefourteenlanguagesout-of-vocabulary
0
0 comments X p. Extension
pith:TZLBPWQO Add to your LaTeX paper What is a Pith Number?
\usepackage{pith}
\pithnumber{TZLBPWQO}

Prints a linked pith:TZLBPWQO badge after your title and writes the identifier into PDF metadata. Compiles on arXiv with no extra files. Learn more

read the original abstract

Most statistical machine translation systems cannot translate words that are unseen in the training data. However, humans can translate many classes of out-of-vocabulary (OOV) words (e.g., novel morphological variants, misspellings, and compounds) without context by using orthographic clues. Following this observation, we describe and evaluate several general methods for OOV translation that use only subword information. We pose the OOV translation problem as a standalone task and intrinsically evaluate our approaches on fourteen typologically diverse languages across varying resource levels. Adding OOV translators to a statistical machine translation system yields consistent BLEU gains (0.5 points on average, and up to 2.0) for all fourteen languages, especially in low-resource scenarios.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.