REVIEW 6 cited by
The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.
Forward citations
Cited by 6 Pith papers
-
ARC-Encoder: learning compressed text representations for large language models
ARC-Encoder pools queries in an encoder's last attention layer to produce compressed continuous representations that a frozen decoder consumes as token embeddings.
-
Building a Functional Machine Translation Corpus for Kpelle
The paper introduces the first claimed public English-Kpelle parallel corpus and shows that fine-tuning NLLB on it yields BLEU up to 30 for Kpelle-to-English.
-
An Empirical Study of Many-to-Many Summarization with Large Language Models
Instruction-tuned open LLMs beat zero-shot GPT-4 on automatic many-to-many summarization scores without lowering MMLU, but human evaluation shows instruction tuning can increase factual errors.
-
Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction
A DQN-based candidate selection and a competitive correction block improve MT ensembling quality while reducing inference cost on English-Hindi and Hindi-English tasks.
-
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
A claimed 2B-parameter multi-task translation model for 36 Indian languages, built from pivoted and synthetic corpora, evaluated without baselines and with inconsistent reported numbers.
-
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
Aya Expanse 8B and 32B report state-of-the-art multilingual win-rates on a new 23-language translated Arena-Hard benchmark, with the 32B beating Llama 3.1 70B by 54.0%.
Discussion (0). Continue with ORCID to comment.