Pith. sign in

REVIEW 4 major objections 4 minor 4 references

The paper claims that showing three machine-translation outputs side by side cuts annotation time by about a third and lowers annotator noise, while preserving absolute quality scores.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 11:44 UTC pith:IDBBZRMS

load-bearing objection The protocol is worth knowing about, but the headline 31% time saving compares cESA to itself, not to standard ESA; the only direct ESA comparison shows cESA is slower and slightly less stable. the 4 major comments →

arxiv 2607.26640 v1 pith:IDBBZRMS submitted 2026-07-29 cs.CL cs.HC

Contrastive ESA: Human Evaluation of Multiple Translations at Once

classification cs.CL cs.HC
keywords contrastive evaluationhuman evaluation of machine translationerror span annotationannotation protocolannotation timeannotator agreementranking stabilitytranslation quality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Human evaluation of machine translation is expensive and noisy. The paper tries to establish that a simple presentation change—showing several translations of the same source document on one screen—makes evaluation both cheaper and more reliable. In a large English-to-Japanese study with professional annotators, the proposed contrastive protocol with three outputs per screen reduced average annotation time per item from 77.8 seconds to 53.7 seconds (roughly 31 percent), while also improving inter-annotator agreement and ranking stability. Even when only one output was shown, the protocol's updated guidelines, tutorial, and anchored 5-percent-step scoring scale reduced annotator disagreement compared with the earlier pointwise error-span protocol. If correct, the work gives evaluation practitioners a ready-to-use protocol that yields stable, interpretable model rankings from simple score averages.

Core claim

Contrastive Error Span Annotation (cESA) shows several translations of the same source document side by side; annotators mark minor and major error spans and assign each output an absolute 0–100 score on an anchored 5%-step scale. The paper's central claim, tested on English→Japanese translations of twelve systems with professional annotators, is that k=3 outputs per screen is the sweet spot: it cuts average per-item annotation time from 77.8s to 53.7s (about 31%) while lowering inter-annotator disagreement and improving ranking stability relative to k=1. A separate ablation shows that even at k=1, the updated guidelines, tutorial, and anchored scale improve annotation quality over the earli

What carries the argument

The central object is the cESA annotation screen: k outputs from different systems (k=1..4) arranged in columns beside the same source document, with error-span highlighting and a single absolute score slider per output. The mechanism that makes it work is shared context—annotators read the source once, spot an error in one translation and then look for the same kind of error in others, and calibrate scores within the screen (the 'joint vs separate evaluation' effect). Work it is doing: turning annotation cost from a per-output reading into a per-screen reading, and producing directly comparable absolute scores.

Load-bearing premise

The 'higher-quality annotation' claim rests on the assumption that inter-/intra-annotator agreement and ranking stability are the right measures of annotation quality, measured on a single language pair (English→Japanese) with one pool of professional annotators; if those metrics do not track true quality, or the results don't generalize, the cost-saving claim survives but the quality claim does not.

What would settle it

Re-run the same comparison on a second language direction (for example, German→English) with a separate pool of professional annotators, and check whether k=3 still beats k=1 on inter-annotator mean absolute error and ranking stability while remaining fastest; also compare both protocols against a small gold set of expert-verified error annotations to test whether agreement gains reflect true accuracy. If the k=3 advantage disappears or the agreement metrics diverge from the gold standard, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Adopting k=3 cESA instead of pointwise ESA cuts annotation time by about 31% per item, directly lowering the cost of large-scale evaluation campaigns.
  • Average cESA scores can be used as-is to rank and compare systems; no TrueSkill-style latent model or post-hoc normalization is needed.
  • Even without showing multiple outputs, the protocol's tutorial, anchored 5%-step scale, and refined guidelines reduce annotator noise.
  • The sweet spot at k=3 suggests a practical ceiling: showing four outputs was slower and less stable, so the optimal tradeoff is not 'more is better'.
  • The protocol's small decoy and similarity biases are negligible enough that random screen assignment plus score averaging preserves ranking validity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the k=1 improvement came from interface and guideline changes alone, the contrastive display and the protocol's quality gains are separable; a future campaign could adopt the guidelines without side-by-side display, or vice versa, to isolate their effects.
  • The strong time saving at k=2 and flat pattern thereafter hints that the main cost driver is repeated source reading; if so, similar savings should appear for other language pairs and modalities, a testable prediction.
  • The small decoy effect (−0.5 points when a stronger system is on screen) implies that random screen assignment is important; practitioners who group systems by quality to save effort may bias scores.
  • If cESA's 5%-step anchored scale becomes standard, scores may become interpretable across campaigns, enabling meta-analyses that current noisy continuous scales do not support.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces Contrastive Error Span Annotation (cESA), a human evaluation protocol in which annotators see k machine-translation outputs for the same document side by side, mark error spans, and assign a 0-100 score. It reports two campaigns: a small-scale ESA-vs-cESA(k=1) comparison and a large-scale cESA campaign with k=1,2,3,4, together with analyses of position, similarity, and surrounding-quality biases. The paper claims that cESA with k=3 saves 31% of annotation time relative to standard pointwise ESA while increasing annotation quality, and recommends k=3 for production use.

Significance. If the central claim were supported, cESA would be a useful, practical contribution: it ships with a ready-to-use implementation, detailed guidelines and tutorials, and a direct comparison of display multiplicity (k) that is rare in the human-evaluation literature. The bias analyses in Section 4.2 are thoughtful and the weak effects are reassuring. These strengths do not compensate for the fact that the headline claim—reductions in time and noise compared to standard pointwise evaluation—is not actually established by the reported experiments.

major comments (4)
  1. [§1 third paragraph; Table 1] The 31% savings claim compares cESA(k=1) at 77.8s with cESA(k=3) at 53.7s, both in the bottom (large-scale) block of Table 1. The only controlled ESA-vs-cESA comparison, in the top block, shows cESA(k=1) is slower than ESA (167.3s vs 151.0s per segment). The two campaigns differ in annotator pool, model set, and protocol details, so the 31% figure is not evidence about standard pointwise ESA; no ESA arm exists in the large-scale campaign.
  2. [§4.1 'Main result: cESA (k=1) is higher quality than ESA'; Table 1 top] The text states that cESA(k=1) yields 'lower inter- and intra-annotator disagreement and higher stability' relative to ESA. Table 1 shows InterAA MAE 13.4 vs 15.9 and IntraAA MAE 5.5 vs 5.9 (improvements), but stability is 0.962 vs 0.964 (a worsening). The claim of higher quality is therefore metric-dependent and internally contradicted. No external criterion is provided to adjudicate between the conflicting metrics, so the abstract's 'reductions in ... noise' is only partially supported.
  3. [§4.1 'Main result: Annotating three models at the same time is faster'; Table 1 bottom] Within the large-scale cESA campaign, k=3 is not uniformly the optimum: it has the lowest InterAA MAE and highest stability, but its IntraAA MAE (10.5) is worse than k=1 (9.7) and k=4 (8.4), and the stability differences across k (0.786-0.813) are small. Without pairwise significance tests or an external validation of the quality measures, the recommendation of k=3 is only weakly supported. More importantly, none of these rows involve standard ESA, so they cannot support the comparative claim against pointwise evaluation.
  4. [Abstract and §5] The abstract claims 'reductions in annotation time and noise compared to standard pointwise evaluation' and §5 claims 'higher-quality annotations than available alternatives.' Given the confounded time comparison and the metric disagreement in the direct ESA comparison, these claims overstate the evidence. The paper should either present a same-campaign ESA-vs-cESA(k=3) comparison or substantially narrow the claims to within-cESA comparisons across k.
minor comments (4)
  1. [§4] The design is described as 'within-subject,' but the text then states that each annotator saw a document in only one condition. This appears to be a between-subject manipulation of conditions per document; clarify the experimental design and the random assignment procedure.
  2. [Table 1 caption/footnote] The footnote refers to 'Table 2 second column, bottom part' where it should refer to Table 1. Also, 'esimate' is a typo in §4.
  3. [§4.2, Tables 5 and 6] The effect sizes in Tables 5 and 6 are reported as raw score differences without confidence intervals or significance tests. Since the conclusion is that biases are negligible, please report uncertainty (e.g., bootstrap intervals) or at least state the standard error.
  4. [Appendix A, Listings 3-4] The two annotation guideline listings are similar but not aligned: ESA uses 0/33/66/100 anchors while cESA uses five bands with 5% steps. Make explicit that these differences are intentional and that Listing 4 is the 'previous work' baseline in the small-scale campaign.

Circularity Check

0 steps flagged

No significant circularity; the validation is empirical and self-contained, though the headline ESA comparison is confounded.

full rationale

This paper is an empirical protocol-validation study rather than a derivation, so the circularity patterns do not directly apply. The central claims rest on human-evaluation data in Table 1: the 31% time saving is computed as 77.8s (cESA k=1, large campaign) to 53.7s (cESA k=3), i.e., within the proposed protocol, not against the standard ESA row; footnote 4 explicitly attributes the small- versus large-campaign cESA(k=1) difference to crowd/model selection. That is a comparison-validity or reporting concern about the phrase 'compared to standard pointwise evaluation,' not a reduction of the conclusion to its inputs by construction. The inter-/intra-annotator MAE and stability measures are operational definitions of annotation quality, but they are standard reliability metrics applied to collected annotations and are not defined in terms of the paper's conclusions; no fitted parameter is renamed as a prediction. The lack of external validation of 'annotation quality' is a validity limitation, not circularity. Self-citations to prior ESA (Kocmi et al., 2024) and Pearmut (Zouhar & Kocmi, 2026) are baseline/implementation references and are not load-bearing uniqueness arguments. No equation, fitted quantity, or definitionally linked pair of results is equivalent to the claimed outcome by construction, so no circularity is demonstrable.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The central claim rests on internal metrics rather than an external benchmark; the protocol itself is the object of study.

axioms (4)
  • domain assumption Annotation quality is adequately captured by inter-/intra-annotator agreement and ranking stability
    Section 4.1 defines quality only through internal consistency metrics; no external gold standard is used.
  • domain assumption Results generalize from En->Ja and the specific professional annotator pool to other language pairs and crowds
    Section 4 reports only English->Japanese; generalization is assumed.
  • domain assumption Random shuffling of model positions suffices to control position bias
    Section 4.2 relies on shuffling to mitigate bias; no formal control.
  • domain assumption The anchors in Listing 2 are interpreted consistently by annotators
    Section 3 defines anchors; no validation of interpretation.

pith-pipeline@v1.3.0-daily-deepseek · 12490 in / 15182 out tokens · 129798 ms · 2026-08-01T11:44:24.762394+00:00 · methodology

0 comments
read the original abstract

Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.

Figures

Figures reproduced from arXiv: 2607.26640 by Marine Carpuat, Martin Popel, Parker Riley, Philipp Koehn, Rachel Bawden, Roman Grundkiewicz, Sara Rajaee, Tom Kocmi, Vil\'em Zouhar.

Figure 1
Figure 1. Figure 1: Comparison of pointwise and contrastive ESA. In the standard pointwise ESA (top), some errors go undetected. In contrastive ESA (bottom), multiple outputs are shown at the same time, which also speeds up the annotation. shown next to a single translation. In contrast, some approaches show multiple outputs at the same time, but either in a scenario where all model outputs can be shown next to each other (Bo… view at source ↗
Figure 2
Figure 2. Figure 2: Two screenshots of document-level contrastive ESA (cESA) with four model outputs shown next to each other. Each segment translation has a marked list of error spans minor or major and final score from 0 to 100%. See more screenshots and annotation guidelines in Appendix A and Appendix B. MQM, which yielded higher-quality annotations. However, this approach does not clearly scale up beyond two document outp… view at source ↗
Figure 3
Figure 3. Figure 3: Screenshots of four configurations of contrastive ESA in Pearmut. The widths of the model output panels [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Sequence of six introductory tutorial steps that teaches the annotators the translation evaluation task. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

4 extracted references · 2 linked inside Pith

  1. [5]

    In Proceedings of the Tenth Conference on Machine Translation, pages 887–904, Suzhou, China

    COMET­poly: Machine Translation Metric Grounded in Other Candidates . In Proceedings of the Tenth Conference on Machine Translation, pages 887–904, Suzhou, China. A cESA and ESA Annotation Guidelines The full annotation guidelines are shown in Listing 3 and Listing 4. They are included by default when using the Pearmut tool ( Listing 1). The annotation gu...

  2. [2018]

    RankME: Reliable Human Ratings for Natural Language Generation . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers) , pages 72–78, New Orleans, Louisiana. Martin Popel, Marketa Tomkova, Jakub Tomek, Łukasz Kaiser, Jakob Uszkoreit, Ondřej...

  3. [2024]

    In Proceedings of the Ninth Conference on Machine Translation, pages 1440–1453, Miami, Florida, USA

    Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation . In Proceedings of the Ninth Conference on Machine Translation, pages 1440–1453, Miami, Florida, USA. Chris Lee, Albert Gatt, Emiel Miltenburg, Sander Wubben, and Emiel Krahmer. 2019. Best practices for the human evaluation of automatically generated text. In Proceedin...

  4. [2025]

    AI­Assisted Human Evaluation of Machine Translation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages 4936–4950, Albuquerque, New Mexico. Maike Züfle, Vilém Zouhar, Tu Anh Dinh, Felipe Maia Polo, Jan Niehues, and Mrinma...