REVIEW 2 cited by
The IIT Bombay English-Hindi Parallel Corpus
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present the IIT Bombay English-Hindi Parallel Corpus. The corpus is a compilation of parallel corpora previously available in the public domain as well as new parallel corpora we collected. The corpus contains 1.49 million parallel segments, of which 694k segments were not previously available in the public domain. The corpus has been pre-processed for machine translation, and we report baseline phrase-based SMT and NMT translation results on this corpus. This corpus has been used in two editions of shared tasks at the Workshop on Asian Language Translation (2016 and 2017). The corpus is freely available for non-commercial research. To the best of our knowledge, this is the largest publicly available English-Hindi parallel corpus.
Forward citations
Cited by 2 Pith papers
-
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
Using 25-30 seconds of audio per speaker from 100 rural Bhojpuri women, synthetic speech augmentation cuts ASR word error on the new SRUTI benchmark by 4.7 points.
-
Pivot Language for Low-Resource Machine Translation
Hindi-pivot transfer gives a 14.2 SacreBLEU on Nepali-English devtest, beating the fully supervised direct baseline by 6.6 points, but with no code, no error bars, and a non-controlled baseline.
Discussion (0). Sign in to comment.