A new minimal-pair benchmark for Urdu grammar shows that multilingual models vary widely across syntactic phenomena, with LLaMA-3-70B best at 94.7% accuracy.
In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
A new minimal-pair benchmark for Urdu grammar shows that multilingual models vary widely across syntactic phenomena, with LLaMA-3-70B best at 94.7% accuracy.