Pith. sign in

REVIEW

Low Resourced Machine Translation via Morpho-syntactic Modeling: The Case of Dialectal Arabic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1712.06273 v1 pith:WDPEK5TN submitted 2017-12-18 cs.CL

classification cs.CL
keywords arabicdatamodelingparalleltranslationdialect-to-dialectdialectalmachine
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the second ever evaluated Arabic dialect-to-dialect machine translation effort, and the first to leverage external resources beyond a small parallel corpus. The subject has not previously received serious attention due to lack of naturally occurring parallel data; yet its importance is evidenced by dialectal Arabic's wide usage and breadth of inter-dialect variation, comparable to that of Romance languages. Our results suggest that modeling morphology and syntax significantly improves dialect-to-dialect translation, though optimizing such data-sparse models requires consideration of the linguistic differences between dialects and the nature of available data and resources. On a single-reference blind test set where untranslated input scores 6.5 BLEU and a model trained only on parallel data reaches 14.6, pivot techniques and morphosyntactic modeling significantly improve performance to 17.5.

Discussion (0). Sign in to comment.

Pith tools