Pith. sign in

REVIEW 3 cited by

The FLoRes Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.01382 v3 pith:XWHF5ZQI submitted 2019-02-04 cs.CL

classification cs.CL
keywords availabledatalow-resourcefloresbecausedatasetsevaluationexperiments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For machine translation, a vast majority of language pairs in the world are considered low-resource because they have little parallel data available. Besides the technical challenges of learning with limited supervision, it is difficult to evaluate methods trained on low-resource language pairs because of the lack of freely and publicly available benchmarks. In this work, we introduce the FLoRes evaluation datasets for Nepali-English and Sinhala-English, based on sentences translated from Wikipedia. Compared to English, these are languages with very different morphology and syntax, for which little out-of-domain parallel data is available and for which relatively large amounts of monolingual data are freely available. We describe our process to collect and cross-check the quality of translations, and we report baseline performance using several learning settings: fully supervised, weakly supervised, semi-supervised, and fully unsupervised. Our experiments demonstrate that current state-of-the-art methods perform rather poorly on this benchmark, posing a challenge to the research community working on low-resource MT. Data and code to reproduce our experiments are available at https://github.com/facebookresearch/flores.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling Language Modeling and Boundaries

    cs.CL 2026-08 reject novelty 6.0 of 10

    The paper hypothesizes that next-byte and boundary distributions in byte-level LMs can be disentangled, proposes two experiments to test it, but provides no experimental results.

  2. From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    English-style debiasing methods only partially transfer to Maltese language models, with Counterfactual Data Augmentation most effective but hampered by grammatical errors.

  3. Pivot Language for Low-Resource Machine Translation

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Hindi-pivot transfer gives a 14.2 SacreBLEU on Nepali-English devtest, beating the fully supervised direct baseline by 6.6 points, but with no code, no error bars, and a non-controlled baseline.

Pith tools