Pith. sign in

REVIEW 3 major objections 5 minor 53 references

PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation

T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read The paper claims that reference-free MT evaluation is better cast as a graded pairwise comparison than as single-candidate scoring, and that this lets smaller models beat far larger metrics on WMT24.

desk verdict Solid pairwise QE paper with a clean controlled comparison; the central claim survives the stress-test better than expected, but the WMT24 numbers need error bars and a configuration lock. read the letter →

arxiv 2601.18006 v2 pith:M5RGJLEY submitted 2026-01-25 cs.CL

classification cs.CL
keywords machinetranslationevaluationqualityestimationpairwisecomparisonrelativescoringgradedpreferencemeta-evaluationminimumBayesriskdecodingantisymmetry
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Modern machine translation evaluation mostly scores one candidate translation at a time against a source, an approach this paper argues is mismatched to the comparative way metrics are actually used. PEAR instead grades two candidate translations jointly, outputting a signed scalar that says which is better and by how much, trained on differences of human segment-level judgments with an antisymmetry regularizer. The authors report that on the WMT24 meta-evaluation benchmark this pairwise formulation beats strictly matched single-candidate QE baselines with identical data and backbones, and that a 560M-parameter model surpasses far larger QE metrics and reference-based metrics. The paper also finds PEAR's signal is less redundant with other top metrics, and that its antisymmetry lets minimum-bayes-risk decoding halve pairwise scoring cost with negligible impact. If correct, the paper establishes a more efficient and accurate formulation for the core comparative task in MT evaluation.

What carries the argument

The central mechanism is the pairwise relative-scoring head operating on a joint cross-encoder representation. The source and two candidates are serialized into one sequence; masked mean pooling produces source and candidate span vectors; a shared utility head maps each candidate to a scalar; the predicted quality difference is the scaled subtraction of the two utilities. An auxiliary antisymmetry loss λflip·(Δ_ab + Δ_ba)^2 enforces sign inversion under candidate order reversal, which the paper shows also improves transitivity over system triples and makes a cheap bidirectional inference mode and an MBR cost-halving shortcut possible.

What would settle it

Retrain PEAR and its matched single-candidate baseline with λflip and δ fixed before any WMT24 evaluation (e.g., chosen on WMT23 held-out data only), then compare on a later MQM meta-evaluation such as WMT25; if the pairwise advantage vanishes or reverses, the reported gains are an artifact of benchmark-informed tuning rather than an intrinsic property of the pairwise formulation.

Watch

Extended reading notes

Core claim

The core claim is that predicting a graded relative preference jointly from a source and two candidates is a more effective formulation for comparative MT evaluation than regressing an absolute score per candidate and subtracting. PEAR encodes source and both candidates in a single cross-encoder, pools masked span representations, maps each candidate to a scalar utility through a shared head, and emits the scaled difference as the relative score. A Huber loss fits human pairwise differences, and an antisymmetry regularization term penalizes deviations from sign inversion under candidate order reversal. The paper argues that this formulation isolates the benefit of pairwise QE: across backbon

Load-bearing premise

The claim rests on the WMT24 meta-evaluation being an independent test set, but the paper's Appendix C.2 says the antisymmetry term and hyperparameters (λflip=0.1, δ=4.5) were chosen after observing WMT24 Avg Corr, so the reported advantage over matched baselines is at least partly fitted to the benchmark.

Editorial extensions

If this is right

  • If the central claim holds, supervised MT metrics should be trained as pairwise comparators rather than absolute scorers when the downstream use is comparison, since the pairwise formulation consistently beats matched single-candidate baselines.
  • Reference-free pairwise evaluation at 560M parameters can match or exceed much larger QE models and reference-based metrics, reducing the compute budget needed for reliable automatic evaluation.
  • The reference-anchored inference mode lets evaluators score N systems in N forward passes rather than N(N−1)/2 pairwise passes, with effectiveness preserved even when the anchor is an MT output rather than a human reference.
  • The reduced correlation of PEAR's signal with other top metrics implies it captures a complementary evaluation signal, which can be exploited in ensembles or as a reward for fine-tuning.
  • Because PEAR's antisymmetry holds in practice, MBR decoding can halve its pairwise scoring cost by computing only one triangle of the utility matrix with negligible quality loss.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be to train a pairwise metric that also conditions on a reference, since anchoring already shows gains; the graded pairwise output could then replace absolute-reference metrics in reference-based evaluations.
  • The antisymmetry and transitivity improvements suggest that structural consistency constraints on relative scores may be a general lever for metric quality; other pairwise evaluation tasks could benefit from an explicit sign-inversion regularizer.
  • Because the antisymmetry term and hyperparameters were chosen after observing WMT24 Avg Corr, a strong test is to freeze them and re-evaluate on a later WMT metrics task; if the pairwise advantage persists, the formulation is genuinely superior.
  • The signed magnitude output could support not just ranking systems but also estimating the statistical significance of differences or calibrating preference strengths, extending PEAR beyond current meta-evaluation statistics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces PEAR, a reference-free supervised MT evaluation metric family that replaces single-candidate absolute scoring with graded pairwise comparison: given a source and two candidate translations, it predicts a signed scalar quality difference. PEAR is trained on WMT DA/DA+SQM/MQM segment-level human judgments by subtracting human scores to form pairwise targets, with an auxiliary antisymmetry loss that encourages sign inversion under candidate-order reversal. The experiments compare PEAR against strictly matched single-candidate QE baselines on the WMT24 MQM benchmark, report gains over larger QE and reference-based metrics, analyze segment-level correlation with other metrics, and demonstrate PEAR as a utility for MBR decoding. The central claim is that the pairwise formulation itself, rather than backbone or training data, drives the observed improvements.

Significance. If the central claim is upheld, the paper makes a useful methodological contribution: it provides a controlled comparison showing that a graded pairwise formulation can outperform matched single-candidate QE under the same data and backbone, which is the right way to isolate the contribution of pairwise versus absolute scoring. The release of code and checkpoints is a concrete strength, as is the careful reporting of multiple inference modes and the MBR analysis. However, the strength of the evidence is reduced by two issues. First, Appendix C.2 reveals that the antisymmetry objective and hyperparameters were selected after observing WMT24 Avg Corr, so the WMT24 benchmark is not fully external to model configuration. Second, all headline comparisons are point estimates on three language pairs with no confidence intervals or significance tests, and the reported margins are often below one Avg Corr point. These issues directly affect the paper's main 'isolating the benefit' claim, even though the underlying idea remains plausible and the experimental design is otherwise well conceived.

major comments (3)
  1. [Section 5.1 and Appendix C.2] The paper's central claim is that the pairwise formulation, not the backbone or data, is responsible for the gains in Table 1. Appendix C.2, however, states that 'including the antisymmetry auxiliary training objective yielded a higher Avg Corr on the overall WMT24 evaluation. Without it, PEAR's Avg Corr decreased from 69.4 to 68.9.' This means the specific configuration used in the headline comparison was selected after inspecting the benchmark on which the comparison is made. Since the antisymmetry term is a distinguishing feature of the pairwise architecture, its inclusion on the basis of WMT24 outcome partially confounds the comparison: the 0.7–1.2-point margins in Table 1 may reflect test-set selection rather than the intrinsic benefit of pairwise scoring. The authors should either provide a WMT23-locked or otherwise fully external configuration-selection protocol, or report results
  2. [Tables 1 and 2] The headline comparisons are reported as point estimates without confidence intervals, bootstrap intervals, or significance tests. The benchmark covers only three language pairs (En-De, En-Es, Ja-Zh), and the principal margins are small: in Table 1, PEAR-XL^both,KD is 70.1 vs. 69.4 for Single-QE-XL^KD, and PEAR^both,KD is 70.0 vs. 69.4 for Single-QE^KD; in Table 2, several comparisons are within 0.2 points of each other. Without uncertainty quantification, a 0.7-point difference on three language pairs could easily be within sampling noise. The authors should report per-language-pair results, bootstrap confidence intervals, or a paired significance test across systems/language pairs for the matched single-candidate comparisons. This is necessary before the claim that PEAR 'outperforms' matched baselines and 'surpasses' larger models can be assessed.
  3. [Section 3.2 and Appendix C.1] The pairwise supervision is constructed by direct subtraction of human segment-level scores: Δ*_ab = s_h(s, mta) − s_h(s, mtb). Appendix C.1 says 'We do not apply any transformation to the scores provided by human annotators.' However, the training data combines DA, DA+SQM, and MQM annotations, which use different scales and protocols (e.g., 0–100 DA scores vs. negative MQM penalties). Directly subtracting scores across these annotation regimes creates training targets whose magnitudes are not comparable across datasets and language pairs. Since the model is trained to regress these differences, this could introduce a scaling artifact rather than a true quality difference. The authors should justify the comparability of these scales, apply a per-dataset or per-annotation-type normalization/standardization, or provide an ablation showing that the results are robust to the choice of target
minor comments (5)
  1. [Table 3] The column header 'Utility LP' is unexplained; the 'LP' abbreviation should be defined or removed. Also, the table mixes metrics under XCOMET-XL/CometKiwi-XL/MetricX-XL without saying which checkpoint of each is used; please specify.
  2. [Figure 2] The caption says 'Ranks are computed by Avg Corr (lower is better).' This is confusing: a rank of 1 is usually best. Please clarify that lower rank values correspond to better Avg Corr, or invert the axis if that is what is intended.
  3. [Appendix D] The correlation matrices label models as 'PEAR' and 'PEAR_GPT-4.1_distil' while the text uses 'PEAR^both' and 'PEAR^both,KD'. Please unify the notation so the reader can map the figures to Tables 1 and 2.
  4. [General] There are several typographical issues, e.g., 'withθ' in Table 1 header, 'V olume' in the references, and 'an ablation where we show' in Section 3.4. A careful proofread is recommended.
  5. [Section 5.4] The analysis of reduced correlation is interesting, but the text should explicitly state whether the Pearson correlations are computed over all segment-system pairs or over a specific subset, and whether the differences in correlation are stable across the three language pairs. As written, the claim of 'less redundant signal' rests on three matrices and a few example correlations.

Circularity Check

1 steps flagged · score 3.0 of 10

The pairwise-vs-single margin is partially WMT24-fitted: Appendix C.2 admits the antisymmetry objective was added after observing WMT24 Avg Corr, so the central 'isolating' claim is not fully out-of-sample; otherwise the derivation chain is self-contained.

  1. fitted input called prediction [Appendix C.2 (Hyperparameters); Section 5.1; Tables 1-2]
    "Additionally, in our initial experiments (Section 5.1), including the antisymmetry auxiliary training objective yielded a higher Avg Corr on the overall WMT24 evaluation. Without it, PEAR’s Avg Corr decreased from 69.4 to 68.9."

    The paper's load-bearing claim is that the pairwise formulation itself, not configuration search, drives the gains ('isolating the benefit of the proposed pairwise formulation', Abstract). Appendix C.2 discloses that the pairwise-specific antisymmetry term was included because it improved Avg Corr on WMT24, and the same WMT24 numbers are then reported as the evaluation of PEAR in Tables 1 and 2. The reported pairwise-vs-single margin is therefore partly a test-set-selected maximum rather than a fully out-of-sample prediction: the model configuration was fit to the benchmark used to score it. This is partial rather than total, since PEAR without the term (68.9) still slightly exceeds the matched Single-QE baseline (68.6).

full rationale

Aside from the benchmark-feedback issue, the derivation chain is self-contained. PEAR is trained on human score differences from WMT16-23/IndicMT and evaluated on WMT24 MQM; there is no definitional circularity, because the prediction target is not defined in terms of the benchmark and WMT24 labels are not training labels. The matched single-candidate baselines share the same backbones, data, and hyperparameters, so the architecture comparison is structurally meaningful. The self-citations in the introduction are motivational rather than load-bearing, and no uniqueness theorem or prior-derived ansatz is imported. The one pith-relevant concern is Appendix C.2's admission that the antisymmetry objective was added after observing WMT24 Avg Corr; this weakens the 'isolating' language and makes the margin partially fitted to the evaluation set. That is a correctness/validity risk rather than a derivation-by-construction equivalence, and it could be resolved by a WMT23-locked configuration search. Hence the moderate score of 3.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

PEAR introduces no new physical or ontological entities; the graded relative score is a model output. The central burden is instead empirical: the paper relies on human score differences as ground truth and on WMT24 being an external benchmark, but the benchmark partly influenced the configuration. Free parameters are standard training hyperparameters plus a learned output scale, with benchmark-dependent selection of λflip and δ being the most important item.

free parameters (4)
  • λflip = 0.1
    Antisymmetry regularization weight set after preliminary HPO; paper states values in [0.1, 0.5] showed no large differences (Appendix C.2).
  • δ (Huber threshold) = 4.5
    Set via HPO over [2.0, 8.0], step 0.5 (Appendix C.2); controls the loss transition from quadratic to linear.
  • α (learned output scale) = learned on MQM stage, initial value 1.0
    α = softplus(α_raw) + ε scales the comparison logit; α_raw is frozen in the first training stage and learned in the second (Section 3.3, Appendix C.2).
  • ε (softplus offset) = small positive constant
    Ensures α > 0; chosen by hand, not estimated from data, but technically a free constant.
assumptions (4)
  • domain assumption Human segment-level scores (DA, DA+SQM, MQM) from different campaigns can be numerically differenced to form reliable pairwise supervision targets.
    Section 3.2 defines Δ*_ab = s_h(s, mt_a) - s_h(s, mt_b), and Appendix C.1 says no transformation is applied across annotation schemes. If the human scores are not interval-comparable, the training target is biased.
  • domain assumption The WMT24 MQM meta-evaluation (En-De, En-Es, Ja-Zh) is a valid external measure of metric quality.
    Section 4.1 adopts the official WMT24 meta-evaluation protocol, but Appendix C.2 reveals some training decisions were made using WMT24 Avg Corr, partially violating the external-benchmark assumption.
  • domain assumption A cross-encoder jointly encoding source and two candidates can learn to represent preference direction and magnitude from human score differences.
    Section 3.3 specifies the InfoXLM/XLM-R cross-encoder architecture and pairwise head; no independent evidence is given that this representational choice is sufficient beyond the empirical results.
  • domain assumption Antisymmetry regularization improves evaluation quality without hurting calibration.
    Section 3.4 and Appendix A measure residual antisymmetry and transitivity deviations, but the downstream benefit of the objective is justified mainly by WMT24 Avg Corr, which was also used for model selection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation." pith.science (2026). https://pith.science/paper/M5RGJLEY

@misc{pith2026260118006,
  author       = {Pith},
  title        = {Pith review of: PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M5RGJLEY}},
  note         = {Machine review of arXiv:2601.18006}
}
read the original abstract

We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-free machine translation (MT) evaluation as a graded pairwise comparison. Given a source segment and two candidate translations, PEAR predicts the direction and magnitude of their quality difference. The metrics are trained using pairwise supervision derived from differences in human judgments, with an additional regularization term that encourages sign inversion under candidate order reversal. On the WMT24 meta-evaluation benchmark, PEAR outperforms strictly matched single-candidate QE baselines trained with the same data and backbones, isolating the benefit of the proposed pairwise formulation. Despite using substantially fewer parameters than recent large metrics, PEAR surpasses far larger QE models and reference-based metrics. Our analysis further indicates that PEAR yields a less redundant evaluation signal relative to other top metrics. Finally, we show that PEAR is an effective utility function for minimum Bayes risk (MBR) decoding, reducing pairwise scoring cost at negligible impact.

Figures

Figures reproduced from arXiv: 2601.18006 by the authors.

Figure 1
Figure 1. Segment- and system-level PEAR evalua￾tion. For each source segment si , PEAR compares the system outputs (mtA i , mtB i ) and predicts a rela￾tive score ∆i . The sign indicates which translation is preferred (∆i > 0: mtA i ; ∆i < 0: mtB i ), while the magnitude |∆i | reflects the strength of the prefer￾ence, with smaller magnitudes corresponding to weaker preferences approaching a tie. The system-level PEAR score i… view at source ↗
Figure 2
Figure 2. Rank stability of PEARref across anchors. The leftmost anchor (Human Ref) uses the human reference as the fixed comparison translation, matching [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. En-De Pearson correlation matrix between segment-level pairwise difference scores derived from WMT24 [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: En-Es Pearson correlation matrix between segment-level pairwise difference scores derived from WMT24 [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Ja-Zh Pearson correlation matrix between segment-level pairwise difference scores derived from WMT24 [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 7 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Ond r ej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aur \'e lie N \'e v \'e ol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. ...

  4. [4]

    Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2008. https://aclanthology.org/W08-0309/ Further Meta-Evaluation of Machine Translation . in Proceedings of the Third Workshop on Statistical Machine Translation, pages 70--106, Columbus, Ohio. Association for Computational Linguistics

  5. [5]

    Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021. https://doi.org/10.18653/v1/2021.naacl-main.280 I nfo XLM : An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training . in Proceedings of the 2021 Conference of the North American Chapter of the Associatio...

  6. [6]

    Daniel Deutsch, George Foster, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.798 Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration . in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12914--12929, Singapore. Association for Computational Linguistics

  7. [7]

    Kevin Duh. 2008. https://aclanthology.org/W08-0331/ Ranking vs. Regression in Machine Translation Evaluation . in Proceedings of the Third Workshop on Statistical Machine Translation, pages 191--194, Columbus, Ohio. Association for Computational Linguistics

  8. [8]

    Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. https://doi.org/10.1162/tacl_a_00437 Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation . Transactions of the Association for Computational Linguistics, 9:1460--1474

Show all 53 references
  1. [9]

    Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022. https://doi.org/10.1162/tacl_a_00491 High Quality Rather than High Model Probability: Minimum B ayes Risk Decoding with Neural Metrics . Transactions of the Association for Computational Linguistics, 10:811--825

  2. [10]

    Markus Freitag, Nitika Mathur, Daniel Deutsch, Chi-Kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Frederic Blain, Tom Kocmi, Jiayi Wang, David Ifeoluwa Adelani, Marianna Buchicchio, Chrysoula Zerva, and Alon Lavie. 2024. https://doi.org/10.18653/v1/2024.wmt-1.2 Ar...

  3. [11]

    Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021. https://doi.org/10.18653/v1/2021.repl4nlp-1.4 Larger-Scale Transformers for Multilingual Masked Language Modeling . in Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2...

  4. [12]

    Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F

    Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F. T. Martins. 2024. https://doi.org/10.1162/tacl_a_00683 x COMET : Transparent Machine Translation Evaluation through Fine-grained Error Detection . Transactions of the Association for ...

  5. [13]

    Francisco Guzm \'a n, Shafiq Joty, Llu \'i s M \`a rquez, Alessandro Moschitti, Preslav Nakov, and Massimo Nicosia. 2014. https://doi.org/10.3115/v1/D14-1027 Learning to Differentiate Better from Worse Translations . in Proceedings of the 2014 Conference on Empirical Methods i...

  6. [14]

    Francisco Guzm \'a n, Shafiq Joty, Llu \'i s M \`a rquez, and Preslav Nakov. 2015. https://doi.org/10.3115/v1/P15-1078 Pairwise Neural Machine Translation Evaluation . in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th Intern...

  7. [15]

    Dan Hendrycks and Kevin Gimpel. 2016. Gaussian Error Linear Units (GELUs). arXiv preprint arXiv:1606.08415

  8. [16]

    Peter J. Huber. 1964. https://doi.org/10.1214/aoms/1177703732 Robust Estimation of a Location Parameter . The Annals of Mathematical Statistics, 35(1):73--101

  9. [17]

    Marcin Junczys-Dowmunt. 2025. https://doi.org/10.18653/v1/2025.wmt-1.67 GEMBA V2: Ten Judgments Are Better Than One . in Proceedings of the Tenth Conference on Machine Translation, pages 926--933, Suzhou, China. Association for Computational Linguistics

  10. [18]

    Juraj Juraska, Daniel Deutsch, Mara Finkelstein, and Markus Freitag. 2024. https://doi.org/10.18653/v1/2024.wmt-1.35 M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task . in Proceedings of the Ninth Conference on Machine Translation, pages 492--504, Miami...

  11. [19]

    Juraj Juraska, Tobias Domhan, Mara Finkelstein, Tetsuji Nakagawa, Geza Kovacs, Daniel Deutsch, Pidong Wang, and Markus Freitag. 2025. https://doi.org/10.18653/v1/2025.wmt-1.70 M etric X -25 and G em S pan E val: G oogle T ranslate Submissions to the WMT 25 Evaluation Shared Ta...

  12. [20]

    Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 M etric X -23: The G oogle Submission to the WMT 2023 Metrics Shared Task . in Proceedings of the Eighth Conference on Machin...

  13. [21]

    Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.97 The Perils of Using M echanical T urk to Evaluate Open-Ended Text Generation . in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, page...

  14. [22]

    Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...

  15. [23]

    Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Ma...

  16. [24]

    Tom Kocmi and Christian Federmann. 2023 a . https://doi.org/10.18653/v1/2023.wmt-1.64 GEMBA - MQM : Detecting Translation Quality Error Spans with GPT -4 . in Proceedings of the Eighth Conference on Machine Translation, pages 768--775, Singapore. Association for Computational ...

  17. [25]

    Tom Kocmi and Christian Federmann. 2023 b . https://aclanthology.org/2023.eamt-1.19/ Large Language Models Are State-of-the-Art Evaluators of Translation Quality . in Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 193--203,...

  18. [26]

    Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. https://aclanthology.org/2021.wmt-1.57/ To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation . in Proceedings of the...

  19. [27]

    Tom Kocmi, Vil \'e m Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popovi \'c , Mrinmaya Sachan, and Mariya Shmatova. 2024 b . https://doi.org/10.18653/v1/2024.wmt-1.131 Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Tra...

  20. [28]

    Tom Kocmi, Vil \'e m Zouhar, Christian Federmann, and Matt Post. 2024 c . https://doi.org/10.18653/v1/2024.acl-long.110 Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies . in Proceedings of the 62nd Annual Meeting of the Association for Computational Lin...

  21. [29]

    Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled Weight Decay Regularization . in International Conference on Learning Representations

  22. [30]

    Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021. https://doi.org/10.18653/v1/2021.acl-long.566 Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers . in Proceedings of the 59th Annual Meeting of the Association for Computational Ling...

  23. [31]

    Ibraheem Muhammad Moosa, Rui Zhang, and Wenpeng Yin. 2024. https://openreview.net/forum?id=Rry1SeSOQL MT -Ranker: Reference-free machine translation evaluation by inter-system ranking . in The Twelfth International Conference on Learning Representations

  24. [32]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a Method for Automatic Evaluation of Machine Translation . in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...

  25. [33]

    Stefano Perrella, Lorenzo Proietti, Pere-Llu \'i s Huguet Cabot, Edoardo Barba, and Roberto Navigli. 2024 a . https://doi.org/10.18653/v1/2024.emnlp-main.1152 Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics . in Proceedings of the 2024 Conference on...

  26. [34]

    Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Edoardo Barba, and Roberto Navigli. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.856 Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In! in Proceedings of the 62nd Annual Meeting of the...

  27. [35]

    Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Niccol \`o Campolungo, and Roberto Navigli. 2022. https://aclanthology.org/2022.wmt-1.51/ M a TES e: Machine Translation Evaluation as a Sequence Tagging Problem . in Proceedings of the Seventh Conference on Machine Tra...

  28. [36]

    Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . in Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395, Lisbon, Portugal. Association for Computational Linguistics

  29. [37]

    Lorenzo Proietti, Stefano Perrella, and Roberto Navigli. 2025 a . https://doi.org/10.18653/v1/2025.acl-short.63 Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress . in Proceedings of the 63rd Annual Meeting of the Associati...

  30. [38]

    Lorenzo Proietti, Stefano Perrella, Vil \'e m Zouhar, Roberto Navigli, and Tom Kocmi. 2025 b . https://doi.org/10.18653/v1/2025.findings-emnlp.1317 Estimating Machine Translation Difficulty . in Findings of the Association for Computational Linguistics: EMNLP 2025, pages 24261...

  31. [39]

    Guerreiro, Jos \'e Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G

    Ricardo Rei, Nuno M. Guerreiro, Jos \'e Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G. C. de Souza, and Andr \'e F. T. Martins. 2023. https://doi.org/10.18653/v1/2023.wmt-1.73 Scaling up C omet K iwi: Unbabel- IST 2023 Submission for the Quality Estimation Sh...

  32. [40]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A Neural Framework for MT Evaluation . in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, ...

  33. [41]

    Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G

    Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and Andr \'e F. T. Martins. 2022. https://aclanthology.org/2022.wmt-1.60/ C omet K iwi: IST -Unb...

  34. [42]

    Khapra, and Raj Dabre

    Ananya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, and Raj Dabre. 2023. https://doi.org/10.18653/v1/2023.acl-long.795 I ndic MT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for I ndian Languages . in Proceedings ...

  35. [43]

    Keisuke Sakaguchi and Benjamin Van Durme. 2018. https://doi.org/10.18653/v1/P18-1020 Efficient Online Scalar Annotation with Bounded Support . in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 208--218, Me...

  36. [44]

    Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 a . https://doi.org/10.18653/v1/2020.acl-main.704 BLEURT : Learning Robust Metrics for Text Generation . in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online. ...

  37. [45]

    Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, and Ankur Parikh. 2020 b . https://aclanthology.org/2020.wmt-1.102/ Learning to Evaluate Translation Beyond E nglish: BLEURT Submissions to the WMT Metrics 2020 Shared Task ....

  38. [46]

    Yixiao Song, Parker Riley, Daniel Deutsch, and Markus Freitag. 2025. https://doi.org/10.18653/v1/2025.acl-long.1002 Enhancing Human Evaluation in Machine Translation with Comparative Judgement . in Proceedings of the 63rd Annual Meeting of the Association for Computational Lin...

  39. [47]

    Shaomu Tan and Christof Monz. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.217 R e M edy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling . in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages...

  40. [48]

    Brian Thompson, Nitika Mathur, Daniel Deutsch, and Huda Khayrallah. 2024. https://doi.org/10.18653/v1/2024.wmt-1.118 Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy . in Proceedings of the Ninth Conference on Machine Trans...

  41. [49]

    Yang Ye, Ming Zhou, and Chin-Yew Lin. 2007. https://aclanthology.org/W07-0736/ Sentence Level Machine Translation Evaluation as a Ranking . in Proceedings of the Second Workshop on Statistical Machine Translation, pages 240--247, Prague, Czech Republic. Association for Computa...

  42. [50]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr BERTScore: Evaluating Text Generation with BERT . in International Conference on Learning Representations

  43. [51]

    Vil \'e m Zouhar, Tom Kocmi, and Mrinmaya Sachan. 2025 a . https://doi.org/10.18653/v1/2025.naacl-long.255 AI -Assisted Human Evaluation of Machine Translation . in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational L...

  44. [52]

    Vilém Zouhar, Peng Cui, and Mrinmaya Sachan. 2025 b . https://arxiv.org/abs/2501.18251 How to Select Datapoints for Efficient Human Evaluation of NLG Models? Preprint, arXiv:2501.18251

  45. [53]

    Maike Z \"u fle, Vil \'e m Zouhar, Tu Anh Dinh, Felipe Maia Polo, Jan Niehues, and Mrinmaya Sachan. 2025. https://doi.org/10.18653/v1/2025.wmt-1.63 COMET -poly: Machine Translation Metric Grounded in Other Candidates . in Proceedings of the Tenth Conference on Machine Translat...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.