LIMBA is a proposed pipeline that combines collection, grammatical tagging, translation, speech, and generative modules to build language models for low-resource languages, with preliminary Sardinian experiments.
Practical Comparable Data Collection for Low-Resource Languages via Images
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We propose a method of curating high-quality comparable training data for low-resource languages with monolingual annotators. Our method involves using a carefully selected set of images as a pivot between the source and target languages by getting captions for such images in both languages independently. Human evaluations on the English-Hindi comparable corpora created with our method show that 81.1% of the pairs are acceptable translations, and only 2.47% of the pairs are not translations at all. We further establish the potential of the dataset collected through our approach by experimenting on two downstream tasks - machine translation and dictionary extraction. All code and data are available at https://github.com/madaan/PML4DC-Comparable-Data-Collection.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models
LIMBA is a proposed pipeline that combines collection, grammatical tagging, translation, speech, and generative modules to build language models for low-resource languages, with preliminary Sardinian experiments.