Language models now exceed 94% accuracy on Brazilian entrance exams, but still lag on mathematics and specialized engineering exams.
Multilingual Performance Biases of Large Language Models in Education
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Large language models (LLMs) are increasingly being adopted in educational settings. These applications expand beyond English, though current LLMs remain primarily English-centric. In this work, we ascertain if their use in education settings in non-English languages is warranted. We evaluated the performance of popular LLMs on four educational tasks: identifying student misconceptions, providing targeted feedback, interactive tutoring, and grading translations in eight languages (Mandarin, Hindi, Arabic, German, Farsi, Telugu, Ukrainian, Czech) in addition to English. We find that the performance on these tasks somewhat corresponds to the amount of language represented in training data, with lower-resource languages having poorer task performance. Although the models perform reasonably well in most languages, the frequent performance drop from English is significant. Thus, we recommend that practitioners first verify that the LLM works well in the target language for their educational task before deployment.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?
Language models now exceed 94% accuracy on Brazilian entrance exams, but still lag on mathematics and specialized engineering exams.