{"id":"d84fda44-adc8-4bd8-8008-e463a7bb4c88","arxiv_id":"2502.02577","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a blind professional A/B evaluation with full document context, Supertext is preferred at document level in three of four language directions, while segment-level preferences are mostly tied.","lead":"This paper compares two commercial translation services, DeepL and Supertext, in a blind professional A/B test with full document context. Segment-level results are mostly tied, while document-level aggregation favors Supertext in three of four language directions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Document-level conclusion rests on majority-vote aggregation with no significance testing; a binomial test against chance is needed and could change the headline claim.","rationale":"The reader's weakest assumption — that majority-vote aggregation of a single rater's segment preferences is an unvalidated proxy for document-level quality — is correct and is the most load-bearing weak point. I agree with the CONDITIONAL verdict because the paper is transparent, releases data, and its central claim could be salvaged by additional analysis. My concrete concern sharpens the reader's point: even taking the authors' own aggregation rule at face value, the reported document-level counts in the strongest direction (de → it-CH, 7/3/10) are not statistically distinguishable from chance under a binomial test (10 of 17 non-tied, p≈0.63). Moreover, two of the three claimed Supertext directions are only presented graphically, with no numerical counts, so the reader cannot verify the claim from the paper alone. The paper also does not report document-level counts for de → en-GB or de → fr-CH in text, despite the abstract claiming three of four directions. The conclusion that the pattern 'suggests superior consistency across longer texts' is thus an interpretive leap over a statistically fragile majority-vote count. The Limitations section openly concedes the absence of significance testing and inter-annotator agreement, and the evaluation was done by Supertext employees, which the paper acknowledges but dismisses. None of this is misconduct; it is a clear case of an under-powered, under-tested descriptive study whose headline inference exceeds what the numbers support. A computational re-analysis of the released data with a simple binomial or sign test would settle whether the direction-level preferences survive; if they do not, the headline claim would need to be softened from 'preference' to 'observed tendency without statistical significance.'","tokens_in":6047,"tokens_out":1704,"duration_ms":15034,"concrete_test":"Recompute the document-level preference counts for each of the four directions from the released GitHub data, then run a two-sided exact binomial test on the non-tied documents per direction (H0: p=0.5) and a paired sign test on segment preferences per document. Also recompute the three reported Supertext document-level preferences after excluding the 9.5% identical segments. If none of the three directions reaches p<0.05, or if the preference flips when identical segments are excluded, the headline claim of a document-level preference for Supertext is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim — that document-level aggregation reveals a preference for Supertext in three of four directions — depends entirely on Section 5.2's rule: a document counts as 'Supertext better' if more segments favor Supertext than DeepL. With only 20 documents per direction, the strongest Supertext result (de → it-CH: 7 DeepL, 3 equal, 10 Supertext) is 10 of 17 non-tied documents, which a two-sided binomial test against p=0.5 does not reject at α=0.05 (p≈0.63). de → en-GB and de → fr-CH are not even reported numerically as document counts, only as figure bars. If the same aggregation is applied to segment-level preferences that are themselves mostly 'equal' and never tested for significance, the headline claim that Supertext is preferred at document level in three directions is not supported by the reported statistics. The authors explicitly state they 'have yet to conduct a systematic qualitative comparison,' yet the conclusion attributes the pattern to 'consistency across longer texts.' An alternative explanation — that document-level preference is just the arithmetic residue of 9.5% identical segments and small per-segment asymmetries — is not ruled out. This is not a claim of bias but a claim of under-evidenced inference: the paper's own Limitations section concedes there is no inter-annotator agreement or significance testing, making the central ranking fragile.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a bilingual A/B evaluation of two commercial machine translation systems, DeepL and Supertext, on 80 documents and 1033 segments across four language directions. Professional translators rated the outputs with full document context, and the authors derive document-level preferences by aggregating segment-level preferences per document. They report that segment-level ratings show no strong preference in most directions, while document-level aggregation favors Supertext in three of four language directions, which they attribute to superior consistency across longer texts. The evaluation data and scripts are released publicly.","tokens_in":6289,"tokens_out":5215,"duration_ms":50851,"significance":"If the document-level aggregate were statistically robust, the paper would be a useful empirical contribution: it demonstrates that the unit of evaluation can change the ranking of two commercial systems, and it provides a reusable dataset for context-aware MT evaluation. The blind design, the use of professional translators, and the public release of data and scripts are concrete strengths. However, as detailed below, the central document-level claim is currently under-supported because the reported counts do not reach statistical significance, no inter-annotator agreement is measured, and the document-level measure is an untested aggregation of segment-level preferences.","major_comments":[{"comment":"The document-level conclusion is not supported by the reported statistics. The only numeric document count given in the text is de→it-CH: 7 DeepL, 3 equal, 10 Supertext. Among the 17 non-tied documents, a two-sided exact binomial test against p=0.5 gives p≈0.63, so this count does not reject chance. Document counts for de→en-GB and de→fr-CH are not reported numerically, only as bars in Figure 2, and no confidence intervals or significance tests are provided. Even the direction with the largest apparent imbalance (en→de-CH: 13 DeepL, 2 equal, 5 Supertext) has a two-sided binomial p≈0.096 over 18 non-tied documents. The abstract's statement that document-level analysis 'reveals a preference for Supertext in three out of four language directions' therefore overstates what the data show.","section":"Section 5.2"},{"comment":"The 'document-level preference' is defined as a majority vote over one rater's segment-level preferences, not as an independent holistic judgment of document quality. Since each document is assigned to a single rater, there is no inter-annotator agreement estimate, and the Limitations section explicitly acknowledges this. The central claim depends on this aggregation rule, but the paper does not validate it against a direct document-level rating or against any measure of cross-segment consistency. Please report per-document counts for all directions, provide evidence on rater reliability (even a small re-rating subset), and justify or validate the majority-vote aggregation.","section":"Sections 4.3 and 5.2"},{"comment":"The attribution of the aggregated pattern to 'consistency across longer texts' is not supported by the evidence. The paper states that it has 'yet to conduct a systematic qualitative comparison' and offers one anecdotal example in Table 2. Because the document-level measure is derived from segment-level preferences, the observed pattern is equally compatible with small per-segment asymmetries or with the distribution of identical segments (9.5% of all segments, rising to 26.1% in the FAQ subset of de→en-GB). The authors should either provide a direct analysis of terminology consistency across segments or explicitly weaken the causal interpretation.","section":"Section 6"},{"comment":"Segment-level claims also lack uncertainty quantification. For example, the en→de-CH segment counts (88 DeepL vs. 57 Supertext among 145 non-tied segments) may be statistically distinguishable, while the claim of 'no strong preference' in the other three directions rests solely on raw counts. Adding exact binomial tests or confidence intervals to the segment-level counts would make the contrast between segment- and document-level results interpretable, and is necessary before the paper can claim that the two units of measurement lead to different conclusions.","section":"Section 5.1"}],"minor_comments":[{"comment":"The reference list contains typographical errors that should be corrected, e.g., 'W A' for 'WA', 'V olk' for 'Volk', and 'V .' for 'V.'.","section":"References"},{"comment":"The caption does not state the tie-handling rule or the denominators for the aggregated document counts, and only one direction is reported numerically in Section 5.2. The numeric counts should be included in the text or in the figure itself.","section":"Figure 2"},{"comment":"Some glyphs in the source text (e.g., '?' and '≡') appear to be rendering artifacts; the example should be checked to ensure it faithfully reproduces the original source text.","section":"Table 2"},{"comment":"The claim that the selected texts are 'unlikely to be contained in the training data' is stated without verification; phrasing this as an assumption rather than a fact would be more accurate.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The study is a vendor-comparison conducted by the company that operates one of the two systems, with all raters employed by that company. This does not by itself invalidate the data, but it raises the bar for statistical rigor and for independent verification. If the authors add the requested tests, report all document counts, and substantially temper the causal attribution, the paper could become a useful descriptive case study; as currently written, the headline claim is not supported by the reported numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a readable, transparent, small-scale A/B comparison, not a deep methodological advance. The genuinely new bit is the head-to-head with full document context between two commercial systems, and the release of the raw data and scripts. That counts for something; the authors do the field a service by showing how a segment-level tie can become a document-level gap, if their aggregation means what they think.\n\nWhat the paper does well: the setup is sensible and honestly reported. They used professional translators, blind A/B, kept system assignments consistent within a document, and published everything. The limitations section admits the big holes; most papers would hide them. The example in Table 2 is a nice concrete illustration of the consistency phenomenon.\n\nWhere it goes soft: the central claim in the abstract—'document-level analysis reveals a preference for Supertext in three out of four language directions'—is not backed by any significance test or confidence interval. For de→it-CH, the strongest case, the count is 7 DeepL, 3 equal, 10 Supertext; a two-sided binomial test against p=0.5 gives p≈0.63, so chance is not ruled out. For two of the four directions, the document counts appear only as figure bars, not numbers, so the reader cannot check the arithmetic. The paper's own Limitations section concedes there is no inter-annotator agreement and no significance testing. That is not a minor footnote; it is a load-bearing gap, because the whole point of the paper is the discrepancy between segment-level and document-level rankings. With 20 documents per direction and a single rater per document, majority voting over segments is a fragile proxy for document-level quality, and the authors do not validate it against independent document-level judgments or a second rater.\n\nOne more thing: the authors work for Supertext. They disclose this and the A/B is randomized and blind, so I don't infer deliberate bias. But the design means the vendor's own employees are the only raters, which is a genuine independence limitation for a comparison meant to inform buyers.\n\nThe stress-test note is right, and the reader's CONDITIONAL verdict is about right. This is not a rejection-level paper; the research question is timely, the data release is real, and the finding, if it survives proper analysis, would matter for MT benchmarking practice. A serious referee should engage, but the authors need to add significance testing, report document-level counts for all four directions, and get external ratings or at least a second rater for a subset. As it stands, the headline sentence overclaims what the statistics can support.","headline":"Small, honest A/B comparison of DeepL and Supertext whose headline document-level claim is not supported by its own statistics; the released data is worth reanalyzing.","tokens_in":6823,"tokens_out":2246,"would_cite":true,"duration_ms":22314,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Document-level ratings flip DeepL–Supertext ranking in three of four language directions.","keywords":["machine translation evaluation","document-level evaluation","human evaluation","A/B testing","LLM-based translation","translation consistency","DeepL","Supertext"],"falsifier":"Re-run the same 80-document evaluation with several independent professional translators per document and ask each for an explicit document-level preference in addition to per-segment ratings; if the majority-aggregated Supertext preference in three of four directions does not reproduce, or does not correlate with the explicit whole-document preference, the paper's headline finding collapses.","tokens_in":5856,"feed_emoji":"⚖️","tokens_out":4244,"duration_ms":37263,"temperature":0.7,"pith_summary":"This paper compares two commercial machine-translation services, DeepL and Supertext, under conditions designed to use each system's full document context: professional translators read entire unsegmented texts and rated sentence-by-sentence which system's translation was better. At the segment level the two systems are near ties in three of four language directions. When per-segment preferences are aggregated per document, Supertext is preferred in three of four directions, which the authors attribute to better consistency across longer texts; the exception is English-to-German (en → de-CH), where DeepL is preferred at both levels. The study matters because real-world translation use is document-level, so segment-only benchmarks may misrank systems.","feed_headline":"Supertext wins document-level MT test in 3 of 4 language pairs","feed_subtitle":"Segment-level scores show no clear winner; reading whole documents flips the verdict.","key_machinery":"The central mechanism is a blind A/B test with document-level presentation and document-level aggregation. Professional translators see the full source text and both translations side-by-side in original order, so context is available; each segment gets a three-way judgment (A better, B better, equal). A document is then scored by majority vote of its segment judgments. This two-step construction — segment judgment under full context, then per-document majority — is what produces the reversal from a segment-level tie to a document-level preference, and it is the object the paper argues should become standard practice.","core_discovery":"The paper's central claim is that the unit of measurement changes the comparison: pooling all segment judgments shows no strong preference, but counting, for each document, which system was preferred on more segments yields a Supertext preference in three of four language directions (de → en-GB, de → fr-CH, de → it-CH) and a DeepL preference in en → de-CH. The authors interpret the document-level result as evidence that Supertext maintains consistency across longer texts, giving an example where DeepL renders the German word Startseite as 'start page', 'home page', and 'Home page' in one document while Supertext stays consistent. The study is presented as a case for context-sensitive evaluation methodology in the LLM era of machine translation.","pith_inferences":["The paper's design cannot distinguish 'better consistency' from 'aggregation artifact': a document with one clearly better segment may outweigh several ties, so a preference-strength-weighted score (for example, MQM severity) could rank the systems differently.","If context-window utilisation is the driving factor, a direct test is to translate the same documents again after forcing sentence-by-sentence segmentation; the prediction is that Supertext's document-level advantage shrinks or disappears.","The en → de-CH exception suggests that target-language variant control (Swiss German versus standard German) may dominate perceived quality; comparing de-CH with de-DE outputs for the same source text could separate consistency from variant handling.","The public release of data and scripts allows re-aggregation with different rules, so the three-of-four result can be checked against a per-document significance test or a per-rater analysis."],"forward_implications":["If document-level aggregation reflects real-world quality, segment-only benchmarking of commercial MT systems can produce misleading rankings for users who translate full documents.","The Supertext preference in three of four directions supports the paper's proposal that smaller LLM-based providers can compete with dominant closed systems, especially on long-text consistency.","A terminology-consistency explanation predicts that Supertext's advantage will be most visible in documents with repeated terms and cross-paragraph references; the Startseite example illustrates the pattern.","For en → de-CH, users should expect DeepL to remain preferable even with full context, possibly because of within-sentence errors or target-variant mixing in Supertext.","Future benchmarking campaigns should let the systems segment the input themselves and evaluate with full document context rather than feeding pre-split sentences."],"supporting_citations":[{"why":"Supplies the case that human-parity claims collapse under document-level evaluation, motivating this paper's full-context design.","marker":"Läubli et al. (2018)"},{"why":"Large-scale study of human evaluation showing segment-level judgments can mislead; cited as the reason A/B tests alone do not assess error severity.","marker":"Freitag et al. (2021)"},{"why":"Argues for escaping the sentence-level paradigm and backs the claim that NMT processes isolated sentences while LLMs use broader context.","marker":"Post and Junczys-Dowmunt (2023)"},{"why":"Shows LLM adaptation for document-level MT and supports the link between context windows and cross-sentence consistency.","marker":"Wu et al. (2024b)"},{"why":"WMT24 findings that automatic metrics can mislead when comparing strong MT systems, motivating human evaluation.","marker":"Kocmi et al. (2024)"},{"why":"Describes A/B testing infrastructure; the paper cites it as the basis for using A/B tests to compare MT systems.","marker":"Tang et al. (2010)"}],"fun_headline_variants":["Document-level test flips verdict: Supertext beats DeepL","Supertext preferred in 3 of 4 language pairs at doc level","Whole-doc context reveals Supertext edge over DeepL","Segment scores tie, but full text favors Supertext","MT comparison flips when translators see full context"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that counting which system was preferred on more segments within a document is a faithful measure of document-level translation quality; the paper does not validate this aggregation against an independent whole-document judgment or report inter-annotator agreement.","fun_headline_variants_meta":{"raw":{"variants":["Document-level test flips verdict: Supertext beats DeepL","Supertext preferred in 3 of 4 language pairs at doc level","Whole-doc context reveals Supertext edge over DeepL","Segment scores tie, but full text favors Supertext","MT comparison flips when translators see full context"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1370,"prompt_tokens":821,"completion_tokens":549,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":467}},"tokens_in":437,"tokens_out":549,"duration_ms":5250,"temperature":1.0,"reasoning_tokens":467,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T11:41:11.199667+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same 80-document evaluation with several independent professional translators per document and ask each for an explicit document-level preference in addition to per-segment ratings; if the majority-aggregated Supertext preference in three of four directions does not reproduce, or does not correlate with the explicit whole-document preference, the paper's headline finding collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes A/B testing infrastructure; the paper cites it as the basis for using A/B tests to compare MT systems."}],"review_version":1}