Pith. sign in

Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese--English news translation task. We empirically test this claim with alternative evaluation protocols, contrasting the evaluation of single sentences and entire documents. In a pairwise ranking experiment, human raters assessing adequacy and fluency show a stronger preference for human over machine translation when evaluating documents as compared to isolated sentences. Our findings emphasise the need to shift towards document-level evaluation as machine translation improves to the degree that errors which are hard or impossible to spot at the sentence-level become decisive in discriminating quality of different translation outputs.

fields

cs.CE 1 cs.CL 1

years

2026 1 2019 1

verdicts

UNVERDICTED 2

representative citing papers

Translationese in Machine Translation Evaluation

cs.CL · 2019-06-24 · unverdicted · novelty 6.0

Translationese in MT test sets biases evaluations, supporting exclusion of reverse-created data, re-evaluation of human-parity claims, and power analysis for reliable significance testing.

citing papers explorer

Showing 2 of 2 citing papers.

  • Translationese in Machine Translation Evaluation cs.CL · 2019-06-24 · unverdicted · none · ref 17 · internal anchor

    Translationese in MT test sets biases evaluations, supporting exclusion of reverse-created data, re-evaluation of human-parity claims, and power analysis for reliable significance testing.

  • An Explainable Approach to Document-level Translation Evaluation with Topic Modeling cs.CE · 2026-04-22 · unverdicted · none · ref 17

    A topic-modeling framework measures document-level thematic consistency in translations by aligning key tokens across languages with a bilingual dictionary and scoring via cosine similarity, providing explainable insights beyond sentence-level metrics.