REVIEW 2 cited by
MTet: Multi-domain Translation for English and Vietnamese
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese research community. Combining with previous works on English-Vietnamese translation, we grow the existing parallel dataset to 6.2M sentence pairs. We also release the first pretrained model EnViT5 for English and Vietnamese languages. Combining both resources, our model significantly outperforms previous state-of-the-art results by up to 2 points in translation BLEU score, while being 1.6 times smaller.
Forward citations
Cited by 2 Pith papers
-
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
VN-MTEB is a new 41-dataset Vietnamese benchmark for text embeddings, built by machine-translating MTEB datasets with embedding-based and LLM-based quality filters.
-
An Efficient Approach for Machine Translation on Low-resource Languages: A Case Study in Vietnamese-Chinese
Fine-tuning mBART with TF-IDF-selected back-translated monolingual sentences raises test BLEU from 38.22 to 38.97 for Vietnamese-to-Chinese and from 35.58 to 38.90 for Chinese-to-Vietnamese.
Discussion (0). Continue with ORCID to comment.