REVIEW 3 major objections 5 minor 53 references
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read The paper claims that reference-free MT evaluation is better cast as a graded pairwise comparison than as single-candidate scoring, and that this lets smaller models beat far larger metrics on WMT24.
desk verdict Solid pairwise QE paper with a clean controlled comparison; the central claim survives the stress-test better than expected, but the WMT24 numbers need error bars and a configuration lock. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the pairwise relative-scoring head operating on a joint cross-encoder representation. The source and two candidates are serialized into one sequence; masked mean pooling produces source and candidate span vectors; a shared utility head maps each candidate to a scalar; the predicted quality difference is the scaled subtraction of the two utilities. An auxiliary antisymmetry loss λflip·(Δ_ab + Δ_ba)^2 enforces sign inversion under candidate order reversal, which the paper shows also improves transitivity over system triples and makes a cheap bidirectional inference mode and an MBR cost-halving shortcut possible.
What would settle it
Retrain PEAR and its matched single-candidate baseline with λflip and δ fixed before any WMT24 evaluation (e.g., chosen on WMT23 held-out data only), then compare on a later MQM meta-evaluation such as WMT25; if the pairwise advantage vanishes or reverses, the reported gains are an artifact of benchmark-informed tuning rather than an intrinsic property of the pairwise formulation.
Extended reading notes
Core claim
The core claim is that predicting a graded relative preference jointly from a source and two candidates is a more effective formulation for comparative MT evaluation than regressing an absolute score per candidate and subtracting. PEAR encodes source and both candidates in a single cross-encoder, pools masked span representations, maps each candidate to a scalar utility through a shared head, and emits the scaled difference as the relative score. A Huber loss fits human pairwise differences, and an antisymmetry regularization term penalizes deviations from sign inversion under candidate order reversal. The paper argues that this formulation isolates the benefit of pairwise QE: across backbon
Load-bearing premise
The claim rests on the WMT24 meta-evaluation being an independent test set, but the paper's Appendix C.2 says the antisymmetry term and hyperparameters (λflip=0.1, δ=4.5) were chosen after observing WMT24 Avg Corr, so the reported advantage over matched baselines is at least partly fitted to the benchmark.
Editorial extensions
If this is right
- If the central claim holds, supervised MT metrics should be trained as pairwise comparators rather than absolute scorers when the downstream use is comparison, since the pairwise formulation consistently beats matched single-candidate baselines.
- Reference-free pairwise evaluation at 560M parameters can match or exceed much larger QE models and reference-based metrics, reducing the compute budget needed for reliable automatic evaluation.
- The reference-anchored inference mode lets evaluators score N systems in N forward passes rather than N(N−1)/2 pairwise passes, with effectiveness preserved even when the anchor is an MT output rather than a human reference.
- The reduced correlation of PEAR's signal with other top metrics implies it captures a complementary evaluation signal, which can be exploited in ensembles or as a reward for fine-tuning.
- Because PEAR's antisymmetry holds in practice, MBR decoding can halve its pairwise scoring cost by computing only one triangle of the utility matrix with negligible quality loss.
Reading between the lines
- A natural extension would be to train a pairwise metric that also conditions on a reference, since anchoring already shows gains; the graded pairwise output could then replace absolute-reference metrics in reference-based evaluations.
- The antisymmetry and transitivity improvements suggest that structural consistency constraints on relative scores may be a general lever for metric quality; other pairwise evaluation tasks could benefit from an explicit sign-inversion regularizer.
- Because the antisymmetry term and hyperparameters were chosen after observing WMT24 Avg Corr, a strong test is to freeze them and re-evaluate on a later WMT metrics task; if the pairwise advantage persists, the formulation is genuinely superior.
- The signed magnitude output could support not just ranking systems but also estimating the statistical significance of differences or calibrating preference strengths, extending PEAR beyond current meta-evaluation statistics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces PEAR, a reference-free supervised MT evaluation metric family that replaces single-candidate absolute scoring with graded pairwise comparison: given a source and two candidate translations, it predicts a signed scalar quality difference. PEAR is trained on WMT DA/DA+SQM/MQM segment-level human judgments by subtracting human scores to form pairwise targets, with an auxiliary antisymmetry loss that encourages sign inversion under candidate-order reversal. The experiments compare PEAR against strictly matched single-candidate QE baselines on the WMT24 MQM benchmark, report gains over larger QE and reference-based metrics, analyze segment-level correlation with other metrics, and demonstrate PEAR as a utility for MBR decoding. The central claim is that the pairwise formulation itself, rather than backbone or training data, drives the observed improvements.
Significance. If the central claim is upheld, the paper makes a useful methodological contribution: it provides a controlled comparison showing that a graded pairwise formulation can outperform matched single-candidate QE under the same data and backbone, which is the right way to isolate the contribution of pairwise versus absolute scoring. The release of code and checkpoints is a concrete strength, as is the careful reporting of multiple inference modes and the MBR analysis. However, the strength of the evidence is reduced by two issues. First, Appendix C.2 reveals that the antisymmetry objective and hyperparameters were selected after observing WMT24 Avg Corr, so the WMT24 benchmark is not fully external to model configuration. Second, all headline comparisons are point estimates on three language pairs with no confidence intervals or significance tests, and the reported margins are often below one Avg Corr point. These issues directly affect the paper's main 'isolating the benefit' claim, even though the underlying idea remains plausible and the experimental design is otherwise well conceived.
major comments (3)
- [Section 5.1 and Appendix C.2] The paper's central claim is that the pairwise formulation, not the backbone or data, is responsible for the gains in Table 1. Appendix C.2, however, states that 'including the antisymmetry auxiliary training objective yielded a higher Avg Corr on the overall WMT24 evaluation. Without it, PEAR's Avg Corr decreased from 69.4 to 68.9.' This means the specific configuration used in the headline comparison was selected after inspecting the benchmark on which the comparison is made. Since the antisymmetry term is a distinguishing feature of the pairwise architecture, its inclusion on the basis of WMT24 outcome partially confounds the comparison: the 0.7–1.2-point margins in Table 1 may reflect test-set selection rather than the intrinsic benefit of pairwise scoring. The authors should either provide a WMT23-locked or otherwise fully external configuration-selection protocol, or report results
- [Tables 1 and 2] The headline comparisons are reported as point estimates without confidence intervals, bootstrap intervals, or significance tests. The benchmark covers only three language pairs (En-De, En-Es, Ja-Zh), and the principal margins are small: in Table 1, PEAR-XL^both,KD is 70.1 vs. 69.4 for Single-QE-XL^KD, and PEAR^both,KD is 70.0 vs. 69.4 for Single-QE^KD; in Table 2, several comparisons are within 0.2 points of each other. Without uncertainty quantification, a 0.7-point difference on three language pairs could easily be within sampling noise. The authors should report per-language-pair results, bootstrap confidence intervals, or a paired significance test across systems/language pairs for the matched single-candidate comparisons. This is necessary before the claim that PEAR 'outperforms' matched baselines and 'surpasses' larger models can be assessed.
- [Section 3.2 and Appendix C.1] The pairwise supervision is constructed by direct subtraction of human segment-level scores: Δ*_ab = s_h(s, mta) − s_h(s, mtb). Appendix C.1 says 'We do not apply any transformation to the scores provided by human annotators.' However, the training data combines DA, DA+SQM, and MQM annotations, which use different scales and protocols (e.g., 0–100 DA scores vs. negative MQM penalties). Directly subtracting scores across these annotation regimes creates training targets whose magnitudes are not comparable across datasets and language pairs. Since the model is trained to regress these differences, this could introduce a scaling artifact rather than a true quality difference. The authors should justify the comparability of these scales, apply a per-dataset or per-annotation-type normalization/standardization, or provide an ablation showing that the results are robust to the choice of target
minor comments (5)
- [Table 3] The column header 'Utility LP' is unexplained; the 'LP' abbreviation should be defined or removed. Also, the table mixes metrics under XCOMET-XL/CometKiwi-XL/MetricX-XL without saying which checkpoint of each is used; please specify.
- [Figure 2] The caption says 'Ranks are computed by Avg Corr (lower is better).' This is confusing: a rank of 1 is usually best. Please clarify that lower rank values correspond to better Avg Corr, or invert the axis if that is what is intended.
- [Appendix D] The correlation matrices label models as 'PEAR' and 'PEAR_GPT-4.1_distil' while the text uses 'PEAR^both' and 'PEAR^both,KD'. Please unify the notation so the reader can map the figures to Tables 1 and 2.
- [General] There are several typographical issues, e.g., 'withθ' in Table 1 header, 'V olume' in the references, and 'an ablation where we show' in Section 3.4. A careful proofread is recommended.
- [Section 5.4] The analysis of reduced correlation is interesting, but the text should explicitly state whether the Pearson correlations are computed over all segment-system pairs or over a specific subset, and whether the differences in correlation are stable across the three language pairs. As written, the claim of 'less redundant signal' rests on three matrices and a few example correlations.
Circularity Check
The pairwise-vs-single margin is partially WMT24-fitted: Appendix C.2 admits the antisymmetry objective was added after observing WMT24 Avg Corr, so the central 'isolating' claim is not fully out-of-sample; otherwise the derivation chain is self-contained.
-
fitted input called prediction
[Appendix C.2 (Hyperparameters); Section 5.1; Tables 1-2]
"Additionally, in our initial experiments (Section 5.1), including the antisymmetry auxiliary training objective yielded a higher Avg Corr on the overall WMT24 evaluation. Without it, PEAR’s Avg Corr decreased from 69.4 to 68.9."
The paper's load-bearing claim is that the pairwise formulation itself, not configuration search, drives the gains ('isolating the benefit of the proposed pairwise formulation', Abstract). Appendix C.2 discloses that the pairwise-specific antisymmetry term was included because it improved Avg Corr on WMT24, and the same WMT24 numbers are then reported as the evaluation of PEAR in Tables 1 and 2. The reported pairwise-vs-single margin is therefore partly a test-set-selected maximum rather than a fully out-of-sample prediction: the model configuration was fit to the benchmark used to score it. This is partial rather than total, since PEAR without the term (68.9) still slightly exceeds the matched Single-QE baseline (68.6).
full rationale
Aside from the benchmark-feedback issue, the derivation chain is self-contained. PEAR is trained on human score differences from WMT16-23/IndicMT and evaluated on WMT24 MQM; there is no definitional circularity, because the prediction target is not defined in terms of the benchmark and WMT24 labels are not training labels. The matched single-candidate baselines share the same backbones, data, and hyperparameters, so the architecture comparison is structurally meaningful. The self-citations in the introduction are motivational rather than load-bearing, and no uniqueness theorem or prior-derived ansatz is imported. The one pith-relevant concern is Appendix C.2's admission that the antisymmetry objective was added after observing WMT24 Avg Corr; this weakens the 'isolating' language and makes the margin partially fitted to the evaluation set. That is a correctness/validity risk rather than a derivation-by-construction equivalence, and it could be resolved by a WMT23-locked configuration search. Hence the moderate score of 3.
Assumptions & free parameters
free parameters (4)
- λflip =
0.1
- δ (Huber threshold) =
4.5
- α (learned output scale) =
learned on MQM stage, initial value 1.0
- ε (softplus offset) =
small positive constant
assumptions (4)
- domain assumption Human segment-level scores (DA, DA+SQM, MQM) from different campaigns can be numerically differenced to form reliable pairwise supervision targets.
- domain assumption The WMT24 MQM meta-evaluation (En-De, En-Es, Ja-Zh) is a valid external measure of metric quality.
- domain assumption A cross-encoder jointly encoding source and two candidates can learn to represent preference direction and magnitude from human score differences.
- domain assumption Antisymmetry regularization improves evaluation quality without hurting calibration.
Cite this review
Pith. "Pith review of PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation." pith.science (2026). https://pith.science/paper/M5RGJLEY
@misc{pith2026260118006,
author = {Pith},
title = {Pith review of: PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation},
year = {2026},
howpublished = {\url{https://pith.science/paper/M5RGJLEY}},
note = {Machine review of arXiv:2601.18006}
}
read the original abstract
We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-free machine translation (MT) evaluation as a graded pairwise comparison. Given a source segment and two candidate translations, PEAR predicts the direction and magnitude of their quality difference. The metrics are trained using pairwise supervision derived from differences in human judgments, with an additional regularization term that encourages sign inversion under candidate order reversal. On the WMT24 meta-evaluation benchmark, PEAR outperforms strictly matched single-candidate QE baselines trained with the same data and backbones, isolating the benefit of the proposed pairwise formulation. Despite using substantially fewer parameters than recent large metrics, PEAR surpasses far larger QE models and reference-based metrics. Our analysis further indicates that PEAR yields a less redundant evaluation signal relative to other top metrics. Finally, we show that PEAR is an effective utility function for minimum Bayes risk (MBR) decoding, reducing pairwise scoring cost at negligible impact.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ond r ej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aur \'e lie N \'e v \'e ol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. ...
-
[4]
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2008. https://aclanthology.org/W08-0309/ Further Meta-Evaluation of Machine Translation . in Proceedings of the Third Workshop on Statistical Machine Translation, pages 70--106, Columbus, Ohio. Association for Computational Linguistics
2008
-
[5]
Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021. https://doi.org/10.18653/v1/2021.naacl-main.280 I nfo XLM : An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training . in Proceedings of the 2021 Conference of the North American Chapter of the Associatio...
-
[6]
Daniel Deutsch, George Foster, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.798 Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration . in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12914--12929, Singapore. Association for Computational Linguistics
-
[7]
Kevin Duh. 2008. https://aclanthology.org/W08-0331/ Ranking vs. Regression in Machine Translation Evaluation . in Proceedings of the Third Workshop on Statistical Machine Translation, pages 191--194, Columbus, Ohio. Association for Computational Linguistics
2008
-
[8]
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. https://doi.org/10.1162/tacl_a_00437 Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation . Transactions of the Association for Computational Linguistics, 9:1460--1474
Show all 53 references
-
[9]
Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022. https://doi.org/10.1162/tacl_a_00491 High Quality Rather than High Model Probability: Minimum B ayes Risk Decoding with Neural Metrics . Transactions of the Association for Computational Linguistics, 10:811--825
2022 doi
-
[10]
Markus Freitag, Nitika Mathur, Daniel Deutsch, Chi-Kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Frederic Blain, Tom Kocmi, Jiayi Wang, David Ifeoluwa Adelani, Marianna Buchicchio, Chrysoula Zerva, and Alon Lavie. 2024. https://doi.org/10.18653/v1/2024.wmt-1.2 Ar...
2024 doi
-
[11]
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021. https://doi.org/10.18653/v1/2021.repl4nlp-1.4 Larger-Scale Transformers for Multilingual Masked Language Modeling . in Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2...
2021 doi
-
[12]
Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F. T. Martins. 2024. https://doi.org/10.1162/tacl_a_00683 x COMET : Transparent Machine Translation Evaluation through Fine-grained Error Detection . Transactions of the Association for ...
2024 doi
-
[13]
Francisco Guzm \'a n, Shafiq Joty, Llu \'i s M \`a rquez, Alessandro Moschitti, Preslav Nakov, and Massimo Nicosia. 2014. https://doi.org/10.3115/v1/D14-1027 Learning to Differentiate Better from Worse Translations . in Proceedings of the 2014 Conference on Empirical Methods i...
2014 doi
-
[14]
Francisco Guzm \'a n, Shafiq Joty, Llu \'i s M \`a rquez, and Preslav Nakov. 2015. https://doi.org/10.3115/v1/P15-1078 Pairwise Neural Machine Translation Evaluation . in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th Intern...
2015 doi
-
[15]
Dan Hendrycks and Kevin Gimpel. 2016. Gaussian Error Linear Units (GELUs). arXiv preprint arXiv:1606.08415
2016 arXiv
-
[16]
Peter J. Huber. 1964. https://doi.org/10.1214/aoms/1177703732 Robust Estimation of a Location Parameter . The Annals of Mathematical Statistics, 35(1):73--101
1964
-
[17]
Marcin Junczys-Dowmunt. 2025. https://doi.org/10.18653/v1/2025.wmt-1.67 GEMBA V2: Ten Judgments Are Better Than One . in Proceedings of the Tenth Conference on Machine Translation, pages 926--933, Suzhou, China. Association for Computational Linguistics
2025 doi
-
[18]
Juraj Juraska, Daniel Deutsch, Mara Finkelstein, and Markus Freitag. 2024. https://doi.org/10.18653/v1/2024.wmt-1.35 M etric X -24: The G oogle Submission to the WMT 2024 Metrics Shared Task . in Proceedings of the Ninth Conference on Machine Translation, pages 492--504, Miami...
2024 doi
-
[19]
Juraj Juraska, Tobias Domhan, Mara Finkelstein, Tetsuji Nakagawa, Geza Kovacs, Daniel Deutsch, Pidong Wang, and Markus Freitag. 2025. https://doi.org/10.18653/v1/2025.wmt-1.70 M etric X -25 and G em S pan E val: G oogle T ranslate Submissions to the WMT 25 Evaluation Shared Ta...
2025 doi
-
[20]
Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 M etric X -23: The G oogle Submission to the WMT 2023 Metrics Shared Task . in Proceedings of the Eighth Conference on Machin...
2023 doi
-
[21]
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.97 The Perils of Using M echanical T urk to Evaluate Open-Ended Text Generation . in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, page...
2021 doi
-
[22]
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...
2024 doi
-
[23]
Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Ma...
2022
-
[24]
Tom Kocmi and Christian Federmann. 2023 a . https://doi.org/10.18653/v1/2023.wmt-1.64 GEMBA - MQM : Detecting Translation Quality Error Spans with GPT -4 . in Proceedings of the Eighth Conference on Machine Translation, pages 768--775, Singapore. Association for Computational ...
2023 doi
-
[25]
Tom Kocmi and Christian Federmann. 2023 b . https://aclanthology.org/2023.eamt-1.19/ Large Language Models Are State-of-the-Art Evaluators of Translation Quality . in Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 193--203,...
2023
-
[26]
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. https://aclanthology.org/2021.wmt-1.57/ To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation . in Proceedings of the...
2021
-
[27]
Tom Kocmi, Vil \'e m Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popovi \'c , Mrinmaya Sachan, and Mariya Shmatova. 2024 b . https://doi.org/10.18653/v1/2024.wmt-1.131 Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Tra...
2024 doi
-
[28]
Tom Kocmi, Vil \'e m Zouhar, Christian Federmann, and Matt Post. 2024 c . https://doi.org/10.18653/v1/2024.acl-long.110 Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies . in Proceedings of the 62nd Annual Meeting of the Association for Computational Lin...
2024 doi
-
[29]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled Weight Decay Regularization . in International Conference on Learning Representations
2019
-
[30]
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021. https://doi.org/10.18653/v1/2021.acl-long.566 Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers . in Proceedings of the 59th Annual Meeting of the Association for Computational Ling...
2021 doi
-
[31]
Ibraheem Muhammad Moosa, Rui Zhang, and Wenpeng Yin. 2024. https://openreview.net/forum?id=Rry1SeSOQL MT -Ranker: Reference-free machine translation evaluation by inter-system ranking . in The Twelfth International Conference on Learning Representations
2024
-
[32]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a Method for Automatic Evaluation of Machine Translation . in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...
2002
-
[33]
Stefano Perrella, Lorenzo Proietti, Pere-Llu \'i s Huguet Cabot, Edoardo Barba, and Roberto Navigli. 2024 a . https://doi.org/10.18653/v1/2024.emnlp-main.1152 Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics . in Proceedings of the 2024 Conference on...
2024 doi
-
[34]
Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Edoardo Barba, and Roberto Navigli. 2024 b . https://doi.org/10.18653/v1/2024.acl-long.856 Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In! in Proceedings of the 62nd Annual Meeting of the...
2024 doi
-
[35]
Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Niccol \`o Campolungo, and Roberto Navigli. 2022. https://aclanthology.org/2022.wmt-1.51/ M a TES e: Machine Translation Evaluation as a Sequence Tagging Problem . in Proceedings of the Seventh Conference on Machine Tra...
2022
-
[36]
Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . in Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395, Lisbon, Portugal. Association for Computational Linguistics
2015 doi
-
[37]
Lorenzo Proietti, Stefano Perrella, and Roberto Navigli. 2025 a . https://doi.org/10.18653/v1/2025.acl-short.63 Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress . in Proceedings of the 63rd Annual Meeting of the Associati...
2025 doi
-
[38]
Lorenzo Proietti, Stefano Perrella, Vil \'e m Zouhar, Roberto Navigli, and Tom Kocmi. 2025 b . https://doi.org/10.18653/v1/2025.findings-emnlp.1317 Estimating Machine Translation Difficulty . in Findings of the Association for Computational Linguistics: EMNLP 2025, pages 24261...
2025 doi
-
[39]
Guerreiro, Jos \'e Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G
Ricardo Rei, Nuno M. Guerreiro, Jos \'e Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G. C. de Souza, and Andr \'e F. T. Martins. 2023. https://doi.org/10.18653/v1/2023.wmt-1.73 Scaling up C omet K iwi: Unbabel- IST 2023 Submission for the Quality Estimation Sh...
2023 doi
-
[40]
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A Neural Framework for MT Evaluation . in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, ...
2020 doi
-
[41]
Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \'e G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and Andr \'e F. T. Martins. 2022. https://aclanthology.org/2022.wmt-1.60/ C omet K iwi: IST -Unb...
2022
-
[42]
Khapra, and Raj Dabre
Ananya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, and Raj Dabre. 2023. https://doi.org/10.18653/v1/2023.acl-long.795 I ndic MT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for I ndian Languages . in Proceedings ...
2023 doi
-
[43]
Keisuke Sakaguchi and Benjamin Van Durme. 2018. https://doi.org/10.18653/v1/P18-1020 Efficient Online Scalar Annotation with Bounded Support . in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 208--218, Me...
2018 doi
-
[44]
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 a . https://doi.org/10.18653/v1/2020.acl-main.704 BLEURT : Learning Robust Metrics for Text Generation . in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online. ...
2020 doi
-
[45]
Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, and Ankur Parikh. 2020 b . https://aclanthology.org/2020.wmt-1.102/ Learning to Evaluate Translation Beyond E nglish: BLEURT Submissions to the WMT Metrics 2020 Shared Task ....
2020
-
[46]
Yixiao Song, Parker Riley, Daniel Deutsch, and Markus Freitag. 2025. https://doi.org/10.18653/v1/2025.acl-long.1002 Enhancing Human Evaluation in Machine Translation with Comparative Judgement . in Proceedings of the 63rd Annual Meeting of the Association for Computational Lin...
2025 doi
-
[47]
Shaomu Tan and Christof Monz. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.217 R e M edy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling . in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages...
2025 doi
-
[48]
Brian Thompson, Nitika Mathur, Daniel Deutsch, and Huda Khayrallah. 2024. https://doi.org/10.18653/v1/2024.wmt-1.118 Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy . in Proceedings of the Ninth Conference on Machine Trans...
2024 doi
-
[49]
Yang Ye, Ming Zhou, and Chin-Yew Lin. 2007. https://aclanthology.org/W07-0736/ Sentence Level Machine Translation Evaluation as a Ranking . in Proceedings of the Second Workshop on Statistical Machine Translation, pages 240--247, Prague, Czech Republic. Association for Computa...
2007
-
[50]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr BERTScore: Evaluating Text Generation with BERT . in International Conference on Learning Representations
2020
-
[51]
Vil \'e m Zouhar, Tom Kocmi, and Mrinmaya Sachan. 2025 a . https://doi.org/10.18653/v1/2025.naacl-long.255 AI -Assisted Human Evaluation of Machine Translation . in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational L...
2025 doi
-
[52]
Vilém Zouhar, Peng Cui, and Mrinmaya Sachan. 2025 b . https://arxiv.org/abs/2501.18251 How to Select Datapoints for Efficient Human Evaluation of NLG Models? Preprint, arXiv:2501.18251
2025 arXiv
-
[53]
Maike Z \"u fle, Vil \'e m Zouhar, Tu Anh Dinh, Felipe Maia Polo, Jan Niehues, and Mrinmaya Sachan. 2025. https://doi.org/10.18653/v1/2025.wmt-1.63 COMET -poly: Machine Translation Metric Grounded in Other Candidates . in Proceedings of the Tenth Conference on Machine Translat...
2025 doi
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.