Y-NQ is a 358-question open-book reading comprehension benchmark for English and Yorùbá, and the paper reports that GPT-4o, o1-mini, and Llama-3.1-8b all perform worse on Yorùbá than on English.
Smith, and Yulia Tsvetkov
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Y-NQ: English-Yor\`ub\'a Evaluation dataset for Open-Book Reading Comprehension and Text Generation
Y-NQ is a 358-question open-book reading comprehension benchmark for English and Yorùbá, and the paper reports that GPT-4o, o1-mini, and Llama-3.1-8b all perform worse on Yorùbá than on English.