LLMs are less reliable at tutoring, feedback, and misconception detection in lower-resource languages, and English prompts usually work as well as or better than prompts in the target language.
AI-Assisted Human Evaluation of Machine Translation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Annually, research teams spend large amounts of money to evaluate the quality of machine translation systems (WMT, inter alia). This is expensive because it requires a lot of expert human labor. In the recently adopted annotation protocol, Error Span Annotation (ESA), annotators mark erroneous parts of the translation and then assign a final score. A lot of the annotator time is spent on scanning the translation for possible errors. In our work, we help the annotators by pre-filling the error annotations with recall-oriented automatic quality estimation. With this AI assistance, we obtain annotations at the same quality level while cutting down the time per span annotation by half (71s/error span $\rightarrow$ 31s/error span). The biggest advantage of the ESA$^\mathrm{AI}$ protocol is an accurate priming of annotators (pre-filled error spans) before they assign the final score. This alleviates a potential automation bias, which we confirm to be low. In our experiments, we find that the annotation budget can be further reduced by almost 25% with filtering of examples that the AI deems to be likely to be correct.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multilingual Performance Biases of Large Language Models in Education
LLMs are less reliable at tutoring, feedback, and misconception detection in lower-resource languages, and English prompts usually work as well as or better than prompts in the target language.