A six-year, open, daily-updated database of over two million unassimilated English borrowings in the Spanish press, with an honest evaluation of its detector's real-world precision.
CALCS 2021 Shared Task: Machine Translation for Code-Switched Data
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
To date, efforts in the code-switching literature have focused for the most part on language identification, POS, NER, and syntactic parsing. In this paper, we address machine translation for code-switched social media data. We create a community shared task. We provide two modalities for participation: supervised and unsupervised. For the supervised setting, participants are challenged to translate English into Hindi-English (Eng-Hinglish) in a single direction. For the unsupervised setting, we provide the following language pairs: English and Spanish-English (Eng-Spanglish), and English and Modern Standard Arabic-Egyptian Arabic (Eng-MSAEA) in both directions. We share insights and challenges in curating the "into" code-switching language evaluation data. Further, we provide baselines for all language pairs in the shared task. The leaderboard for the shared task comprises 12 individual system submissions corresponding to 5 different teams. The best performance achieved is 12.67% BLEU score for English to Hinglish and 25.72% BLEU score for MSAEA to English.
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press
A six-year, open, daily-updated database of over two million unassimilated English borrowings in the Spanish press, with an honest evaluation of its detector's real-world precision.