Automatic MT metrics often rank on par with or above human annotators when both are scored against MQM human judgments, raising doubts about whether progress in MT evaluation can still be measured.
Title resolution pending
1 Pith paper cite this work, alongside 6 external citations. Polarity classification is still indexing.
1
Pith paper citing it
6
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
Automatic MT metrics often rank on par with or above human annotators when both are scored against MQM human judgments, raising doubts about whether progress in MT evaluation can still be measured.