{"id":"362c348f-5e06-44f6-9886-55083961d04c","arxiv_id":"2506.19571","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Automatic MT metrics often rank on par with or above human annotators when both are scored against MQM human judgments, raising doubts about whether progress in MT evaluation can still be measured.","lead":"This paper runs human raters through the same ranking pipeline used for machine translation metrics and finds the metrics often match or beat the raters. The result suggests MT evaluation may have hit the ceiling of its own human yardstick, making future progress hard to measure.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on treating agreement with one MQM-based gold evaluator as the yardstick for human baselines from other protocols; since Appendix F only swaps one single-protocol evaluator for another, the 'metrics surpass humans' result may be an artifact of gold choice rather than…","rationale":"Read in good faith, the paper's empirical result is real: using WMT meta-evaluation measures, SOTA metrics are frequently in the same significance cluster as human baselines, and often numerically above them. The authors also flag the main confounds. The reader's weakest-assumption analysis identifies the right load-bearing point: the comparison only means what the abstract suggests if agreement with the chosen MQM gold is a fair yardstick for all evaluators. I agree with that assessment. The remaining gap is that Appendix F, while a useful robustness check, does not settle the yardstick question: it changes which single evaluator is the gold, but never constructs a multi-rater consensus gold. Given the known acc*eq bias toward continuous scales and the very small 2023 EN→DE intersection, the 'higher than human baselines' part of the claim is not yet cleanly separated from protocol/scale mismatch. This does not justify rejection, because the paper's conclusions are carefully hedged and its contribution is to raise the measurement problem, not to assert parity. The proposed consensus-gold recomputation is a feasible, decisive check. Hence the reader's CONDITIONAL verdict is appropriate and should stand unchanged.","tokens_in":25663,"tokens_out":8956,"duration_ms":104923,"concrete_test":"Recompute the main rankings using a multi-rater consensus gold instead of a single evaluator. For 2023 EN→DE and 2024 EN→ES, aggregate all available human annotations (e.g., mean/median of all MQM raters, and a z-score-combined MQM+ESA gold) and re-run SPA and acc*eq for every metric and every individual human evaluator. If metrics no longer rank at or above the median human baseline, the 'human parity' reading is gold-dependent; if they do, the protocol-mismatch concern is substantially resolved. A complementary check: on the 2024 EN→ES 449 segments, collect 2-3 additional MQM and ESA annotations and form a leave-one-rater-out consensus gold, then compare within-protocol human baselines to metrics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 designates MQM-derived evaluators as ground truth and all other protocols (ESA, pSQM, DA+SQM) as human baselines. Section 4.1 itself asks whether a metric that ranks higher is a better evaluator or merely aligns with the MQM score distribution. Appendix F varies the ground truth among pSQM-1, DA+SQM, MQM-2023-2/3, and ESA, but every alternative is still a single evaluator, not a consensus of multiple raters. The 2024 EN→ES set has exactly one MQM and one ESA evaluator, so the cross-protocol comparison rests entirely on that single gold. Under acc*eq the comparison is further confounded: tie calibration favors continuous-scale evaluators, while human protocols produce discrete scores (Perrella et al. 2024b), which can depress human baselines for scale reasons, not quality reasons. The 2023 EN→DE set, where humans most clearly fall below metrics, contains only 145 segments (footnote 3), so that year's evidence is statistically fragile. None of this makes the results uninteresting, but it means the headline claim has not yet been separated from the choice of a single gold evaluator.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper re-runs the WMT 2024 meta-evaluation protocols on seven test sets from the 2020, 2022, 2023, and 2024 WMT Metrics Shared Tasks, adding human annotators as evaluators alongside automatic metrics. Using a single MQM-based human evaluator as ground truth, the authors compute Soft Pairwise Accuracy (SPA) and Pairwise Accuracy with Tie Calibration (acc*_eq) for both metrics and human baselines from other protocols (ESA, pSQM, DA+SQM). The central claim is that automatic metrics often rank in the same statistical significance cluster as human baselines under SPA and frequently surpass them under acc*_eq, which the paper interprets as suggesting, while explicitly cautioning against, human parity in MT evaluation. The paper also discusses the limits of measuring progress when the human reference no longer separates evaluator types.","tokens_in":25910,"tokens_out":9896,"duration_ms":96805,"significance":"If the empirical result is robust, it has important consequences: current WMT meta-evaluation may no longer be able to distinguish automatic metrics from human raters, and incremental metric improvements become ambiguous. The paper makes a useful contribution by bringing human baselines into the standard meta-evaluation framework and by explicitly raising the interpretation problem. Strengths include the use of official WMT meta-evaluation measures, statistical significance clusters on rankings, and a robustness appendix that varies the ground-truth evaluator; the authors also release their code. However, the central claim is currently conditional on the choice of a single gold evaluator and on the treatment of discrete human scores under acc*_eq, which the paper itself identifies but does not resolve experimentally.","major_comments":[{"comment":"The ground-truth evaluator is always a single human evaluator (MQM in the main results; pSQM-1, DA+SQM, MQM-2023-2/3, or ESA in Appendix F), never a consensus of multiple raters, and the 2024 EN→ES test set has exactly one MQM and one ESA evaluator. As the authors themselves ask in §4.1, a high ranking may simply reflect closer alignment with the score distribution of the chosen gold protocol or rater rather than better evaluation capability; this is exactly the sort of protocol-mismatch artifact that could put human baselines below metrics. To support the headline claim that human baselines are not consistently superior, the paper should add a consensus-based gold (e.g., averaging the multiple MQM evaluators available for 2020, 2022, and 2023) or a leave-one-out analysis across gold evaluators, and show that the relative ordering of metrics and human baselines is stable under those conditions.","section":"§2.2 and Appendix F"},{"comment":"The acc*_eq measure systematically ranks human evaluators far lower than SPA does (e.g., pSQM-2 at rank 9 in 2020 EN→DE, DA+SQM at rank 16 in 2022 EN→DE, DA+SQM at rank 14 in 2023 EN→DE), and this is the main source of the 'metrics surpass humans' evidence. The authors attribute this to Perrella et al. (2024b)'s finding that tie calibration favors continuous-scale evaluators while human protocols produce discrete scores, but they do not quantify or control for this scale mismatch. A control experiment that rounds metric scores to the integer granularity of human protocols, or that applies an identical tie-calibration policy to all evaluators, would show whether the acc*_eq gaps are an artifact of score granularity rather than of evaluation quality; without this control, the acc*_eq results do not by themselves establish that metrics outperform human baselines.","section":"§4, Appendix C.2, Table 2"},{"comment":"The 2023 EN→DE test set contains only 145 segments after filtering, yet it is one of the two test sets where human baselines fall most consistently below automatic metrics under both SPA and acc*_eq. The authors acknowledge this small sample in footnote 3, but the main text does not report whether the Table 6 rankings are stable when the test set is enlarged to the 376 segments used in Appendix F (Tables 9–12, with different gold evaluators). Given that the 2023 EN→DE result is a key piece of evidence for the paper's central claim, the paper should either present the Table 6 rankings on the larger segment set with the same MQM gold, or provide confidence intervals for the rank differences, to demonstrate that the 145-segment fragility does not drive the headline result.","section":"Table 1, footnote 3, Table 6"}],"minor_comments":[{"comment":"The sentence 'Following standard practice in the literature ... we designate evaluators derived from the MQM annotations ... as the ground truth' correctly notes the convention, but the paper should also acknowledge that in 2020 the pSQM protocol was also professional and that the choice of MQM as the gold is a substantive modeling decision rather than purely a default.","section":"§2.2"},{"comment":"The caution about the 145-segment 2023 EN→DE set appears only in a footnote at the end of the Discussion; it should be prominently placed near Table 1 or Table 6, since it directly bears on the main result.","section":"Footnote 3"},{"comment":"The use of 'Acc.' and 'Rank' in the captions is clear, but the superscripting of acc*_eq is inconsistent in several places; please unify the notation.","section":"Table 2 and Appendix E"},{"comment":"The phrase 'human evaluators do not consistently rank higher than automatic metrics' could be made more precise by adding 'in the specific test sets and meta-evaluation measures considered here', to avoid overgeneralization.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"This is a thoughtful and honest paper that exposes a real problem in MT meta-evaluation. My main worry is that the headline empirical result is not yet separated from the gold-evaluator artifact and the acc*_eq scale mismatch; with the additional analyses suggested in the major comments, the paper would be a solid contribution. I would reconsider after a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper does something useful that nobody has done cleanly before—it puts human annotators from several protocols (MQM, ESA, pSQM, DA+SQM) into the same WMT meta-evaluation ranking as automatic metrics, with a disjoint-rater constraint so the human baselines aren't inflated. The main empirical finding, that top metrics often land in the same significance cluster as human evaluators and sometimes above them under acc*_eq, survives a reasonable set of robustness checks. That's real work, and the authors deserve credit for it.\n\nWhat's genuinely new: previous work (Perrella et al. 2024a) compared metrics with a single human protocol. Here the comparison is systematic across WMT 2020-2024 and four protocols, with statistical significance clusters and a full appendix varying which evaluator is used as ground truth. The code is released. The paper also does a good job of talking itself out of overclaiming: the discussion of tie calibration, annotation quality, and the 145-segment 2023 EN->DE set is honest and mostly on target.\n\nThe soft spot is the one the reader flagged: the yardstick problem. The meta-evaluation treats one MQM-derived evaluator as ground truth, and then asks whether other humans and metrics agree with it. Under that protocol, a metric that ranks higher than a human from another protocol might just be closer to MQM's score distribution. Appendix F swaps in alternative single evaluators (pSQM, DA+SQM, another MQM, ESA), and the finding mostly holds—but every alternative is still a single rater or a single protocol's aggregate, not a consensus of multiple independent raters. The 2024 EN->ES set has exactly one MQM and one ESA evaluator, so that year's cross-protocol comparison rests on one person's scores on each side. That's thin. And the acc*_eq confound with discrete human scores is real; the tie-calibration algorithm systematically favors continuous metrics, which is likely part of why humans sink under that measure. The authors acknowledge this but don't fully control it. None of this makes the paper wrong; it makes the strong claim (human parity has been reached) not yet established.\n\nBottom line: for the MT evaluation community this is a serious and useful reference point. It deserves a proper peer review with referees who can dig into the gold-evaluator question and the 2024 single-rater sets. I'd take it to reading group and would cite it.","headline":"A careful empirical study that puts human baselines from multiple protocols into WMT meta-evaluation, with honest limitations; the 'parity' result is suggestive but not yet established.","tokens_in":26463,"tokens_out":2543,"would_cite":true,"duration_ms":25991,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human annotators are not consistently better than automatic metrics in machine-translation evaluation: state-of-the-art metrics often rank on par with or above human baselines.","keywords":["machine translation evaluation","meta-evaluation","human parity","human baselines","inter-annotator agreement","soft pairwise accuracy","tie calibration","WMT shared task"],"falsifier":"Take a test set of translations with injected, expert-verified errors that human raters reliably catch—gender agreement, number inflection, named entities, word-sense ambiguities—and recompute SPA and $acc^*_{eq}$ with human baselines and top metrics; if human evaluators then consistently rank above every metric, the reported parity was an artifact of easy test sets and discrete human scales.","tokens_in":25463,"feed_emoji":"📊","tokens_out":9404,"duration_ms":82900,"temperature":0.7,"pith_summary":"This paper asks whether automatic machine-translation metrics have reached human parity as evaluators, and returns a qualified yes. When human annotators are entered into the same leaderboards used to rank metrics, state-of-the-art metrics often match or exceed human baselines: human evaluators share the top statistical cluster with metrics under system-level Soft Pairwise Accuracy and are frequently surpassed under segment-level tie-calibrated Pairwise Accuracy ($acc^*_{eq}$). The comparison draws on seven WMT test sets with overlapping annotations from several human protocols, and raters are partitioned into disjoint evaluator groups so that human–human agreement is not inflated. The authors then argue that these results are not enough to declare human parity, and that they make further progress in MT evaluation hard to measure.","feed_headline":"Top MT metrics now rank with or above human judges","feed_subtitle":"A human-baseline study of WMT leaderboards finds the human ceiling no longer separates people from metrics.","key_machinery":"The machinery is a human-baseline meta-evaluation setup. Each test set is restricted to the largest subset of segments that can be covered by evaluators built from disjoint sets of raters, found by solving an integer linear program, so that no rater contributes to both the ground truth and a baseline. The resulting human evaluators are then ranked against automatic metrics with Soft Pairwise Accuracy (SPA), which rewards an evaluator for expressing confidence levels close to the ground truth when ranking MT systems, and with tie-calibrated Pairwise Accuracy ($acc^*_{eq}$), which counts segment-level pairwise agreements while calibrating ties per evaluator. The interplay of these two measures carries the argument: human evaluators look near the top under SPA but drop under $acc^*_{eq}$, and the paper attributes the drop to human scores being discrete rather than continuous.","core_discovery":"The central claim is that, in current MT meta-evaluation, the human reference no longer separates people from machines. Treating one MQM-based human evaluator as ground truth and other human protocols as baselines, the authors rank every evaluator—human and automatic—on the same WMT 2024 meta-evaluation measures. Human baselines do not consistently rank higher than automatic metrics; under SPA they often sit in the same significance cluster as top metrics, and under $acc^*_{eq}$ they are frequently outranked. The finding is robust to swapping the ground-truth evaluator, but the authors caution that it does not establish equivalence: a metric that ranks above a human may simply align more closely with the score distribution of the protocol used as gold, and current test sets may be too easy for the gap to be meaningful.","pith_inferences":["If the human reference is saturated, optimizing a metric for the leaderboard becomes equivalent to optimizing it for the specific human raters behind the gold annotations, which rewards fitting the benchmark rather than evaluation quality.","A sharper test of parity would inject expert-verified errors (wrong gender, wrong number, named-entity errors, word-sense ambiguities) into translations and check whether human baselines still lose to metrics on those segments; the paper's data cannot answer this.","Building a consensus ground truth from several protocols instead of a single MQM evaluator, or letting human raters give continuous severity scores, might move human baselines to the other side of the parity line.","The gap between SPA and $acc^*_{eq}$ suggests the human disadvantage is largely a segment-level, scale-granularity phenomenon; a human protocol with finer-grained scores could plausibly reverse the ranking."],"forward_implications":["A metric that ranks above a human baseline in a WMT-style leaderboard no longer demonstrates that it evaluates better than a person; it may only show closer agreement with the MQM protocol used as gold.","Rankings under $acc^*_{eq}$ systematically penalize human raters for producing discrete scores, so cross-protocol comparisons of human and automatic evaluators are confounded by score granularity.","If current test sets are as easy as the fluency-only sentinel metric result suggests, then measured parity may vanish once metrics are tested on adversarial or out-of-domain translations.","The field's ability to track improvement in MT evaluation is at risk: once top metrics sit at the human ceiling, a higher ranking is ambiguous and no longer a clear sign of progress."],"supporting_citations":[{"why":"Supplies the 2020 MQM and pSQM annotations and the multi-rater setup from which the 2020 human evaluators and ground truth are derived.","marker":"Freitag et al. (2021a)"},{"why":"Provides the WMT 2024 meta-evaluation strategies used for the rankings and the MQM gold annotations for 2024 EN→ES.","marker":"Freitag et al. (2024)"},{"why":"Introduces Soft Pairwise Accuracy (SPA), the system-level measure that places human baselines in the top cluster.","marker":"Thompson et al. (2024)"},{"why":"Introduces tie-calibrated Pairwise Accuracy ($acc^*_{eq}$), the segment-level measure under which human baselines are often surpassed.","marker":"Deutsch et al. (2023)"},{"why":"First to rank human and automatic evaluators jointly; the paper extends that joint assessment across test sets and protocols.","marker":"Perrella et al. (2024a)"},{"why":"Defines the Error Span Annotation protocol and provides ESA annotations for 2023 EN→DE.","marker":"Kocmi et al. (2024b)"},{"why":"Provides the DA+SQM annotations used as human baselines in the 2022 test sets.","marker":"Kocmi et al. (2022a)"},{"why":"Releases additional MQM annotations for 2022 and argues for more annotation resources to stabilize meta-evaluation.","marker":"Riley et al. (2024)"},{"why":"Shows that a metric can be explicitly specialized to match the gold raters, grounding the paper's caution about what a higher ranking means.","marker":"Finkelstein et al. (2024)"},{"why":"Documents that fine-tuned metrics fail in unseen domains, supporting the paper's claim that current test sets may be too easy.","marker":"Zouhar et al. (2024)"}],"fun_headline_variants":["MT metrics outrank human judges in meta-eval","Human parity in MT evaluation: metrics match or beat humans","Automatic metrics now rival human evaluators on WMT","Human reference no longer separates MT metrics from people","MT meta-eval: metrics surpass human baselines, with caveats"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that agreement with one chosen MQM-based human evaluator is a valid yardstick for ranking both automatic metrics and human raters who used other protocols; if that premise fails, metrics that outrank humans may just be matching the MQM score distribution rather than evaluating better.","fun_headline_variants_meta":{"raw":{"variants":["MT metrics outrank human judges in meta-eval","Human parity in MT evaluation: metrics match or beat humans","Automatic metrics now rival human evaluators on WMT","Human reference no longer separates MT metrics from people","MT meta-eval: metrics surpass human baselines, with caveats"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000367,"raw_usage":{"total_tokens":1935,"prompt_tokens":869,"completion_tokens":1066,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":985}},"tokens_in":485,"tokens_out":1066,"duration_ms":9945,"temperature":1.0,"reasoning_tokens":985,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:06:29.071887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a test set of translations with injected, expert-verified errors that human raters reliably catch—gender agreement, number inflection, named entities, word-sense ambiguities—and recompute SPA and $acc^*_{eq}$ with human baselines and top metrics; if human evaluators then consistently rank above every metric, the reported parity was an artifact of easy test sets and discrete human scales.","supporting_citations":[],"review_version":1}