{"id":"a5fa9a9a-36e3-44af-b1dd-75888deb07be","arxiv_id":"1908.03043","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new English-Czech test suite for WMT19, manually annotated, shows context-aware NMT systems did not outperform sentence-level systems on discourse connectives, AltLexes, or topic-focus word order.","lead":"A team built a test set of 101 English-Czech documents to measure whether machine translation handles text-level meaning, and manually graded five systems on three discourse features. Both document-aware and sentence-level systems performed similarly, with word order errors dominating.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The word-order metric labels the human reference wrong in 6/7 focus-proper and 3/22 placement cases, yet the authors say the reference is not incorrect; this unvalidated binary rule underlies the component where context-aware systems look worst, so the central negative claim is not yet established.","rationale":"The paper's central empirical claim is that document-level Transformer systems showed no measurable gains on AltLexes, connectives, or subject word order. The most load-bearing part of that claim is the word-order evaluation, because it is the component where the context-aware systems appear worst and because its scoring rule is a strong, theory-derived binary judgment. The reader's weakest assumption correctly identifies the indefinite-article proxy as fragile. My review sharpens this: the paper's own reference rows provide direct internal evidence against the validity of the metric. The human reference, which is explicitly described as not incorrect, receives 'no' marks in 6/7 focus-proper and 3/22 placement judgments. If those marks are not errors, then the binary yes/no labels in Section 7.4 do not measure translation quality; they measure conformity to a preferred FGD word-order template. Under that reading, the comparison between context-aware and sentence-level systems in Section 7.4, and the overall conclusion that context-awareness did not help, is not supported. I am not calling the paper fraudulent or dismissing its useful test suite and annotation effort; the concern is about the validity of one scoring rule, and it is testable by independent acceptability judgments. If the test shows the 'no' items are genuinely disfluent, the concern is resolved and the conditional verdict can stand. If it shows they are acceptable, the word-order system comparison should be discarded and the comparative conclusion substantially weakened. This does not change the reader's CONDITIONAL verdict, which already required stronger evidence, data release, or softened claims.","tokens_in":10150,"tokens_out":15765,"duration_ms":178131,"concrete_test":"Extract all 'no' word-order items from Section 7.4, including the reference's 3/22 placement and 6/7 focus-proper cases, and have at least two independent native Czech-speaking linguists (not the original annotators) rate the Czech outputs for acceptability/naturalness on a 1-5 scale, without revealing the FGD-expected order. If a substantial share (say >30%) of the 'no' items are rated acceptable (4 or 5), the binary metric is over-strict; then re-run the word-order system rankings using only items whose 'no' is independently confirmed as disfluent and check whether the context-aware systems are still worse than the sentence-level systems.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the word-order component (Sections 6.1.1 and 7.4), a Czech output is marked incorrect when a contextually non-bound subject introduced by an English indefinite article is not placed in the post-verbal/focus-proper position demanded by the FGD-based rule. The validity of this binary rule is contradicted by the paper's own reference data: the human reference is marked 'no' in 3/22 placement judgments and in 6/7 focus-proper judgments, and Section 8 explicitly says this does not mean the reference is incorrect. A metric that labels the human reference wrong at these rates is measuring a preferred information-structure ordering, not translation adequacy. This matters because word order is the component where the context-aware systems look worst (Marian 14/5, DocTransf-T2T 13/3 vs CUNI-Transf-2018 14/0 in Section 7.4); if those 'no' labels correspond to acceptable Czech word orders, this component provides no evidence that document-level input failed to help. The central 'no measurable gains' claim therefore rests on an unvalidated scoring rule, not merely on small samples.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a manually annotated test suite for evaluating document-level English-to-Czech NMT at WMT19. The authors select 101 PCEDT documents with PDTB-style discourse annotations and use trained linguist annotators to evaluate five systems plus a human reference on three coherence-related phenomena: alternative lexicalizations of discourse connectives (AltLexes), discourse connectives, and topic-focus articulation realized through subject word order. They report inter-annotator agreement (80% overall), qualitative error analyses across linguistic levels, and quantitative comparisons. The central claim is that the two context-aware systems (CUNI-DocTransf-T2T and CUNI-DocTransf-Marian) did not outperform sentence-level systems on the observed phenomena, and that the reference translation was sometimes judged worse without being incorrect.","tokens_in":10364,"tokens_out":4516,"duration_ms":48370,"significance":"The test suite and manual annotation methodology are valuable resources for a difficult evaluation problem, and the paper is transparent about the annotation procedure, including the reporting of inter-annotator agreement. The qualitative inventory of error types (morphology, lexicon, syntax, semantics, discourse) is useful for system developers. The negative finding on context-aware systems is interesting and falsifiable, but as presented it rests on very small counts and on a word-order scoring rule that is not validated against acceptability; if the rule is confirmed by further adjudication, the result would be a meaningful contribution to the document-level MT debate. The authors deserve credit for explicitly noting that the reference translation being marked as 'worse' does not necessarily mean it is incorrect, although this admission creates a tension with the way the word-order metric is used.","major_comments":[{"comment":"The word-order metric labels the human reference as a 'no' in 3/22 placement judgments and in 6/7 focus-proper judgments, while Section 8 states that this 'does not mean that the reference is incorrect.' This is an internal inconsistency in the validity of the metric: if the rule is a measure of translation adequacy, the reference should almost always satisfy it; if the rule encodes a strong FGD-based stylistic preference, then the 'no' labels for MT systems are not evidence of translation errors. Because the two context-aware systems have the worst 'no' counts in the placement table (CUNI-DocTransf-Marian 14/5 and CUNI-DocTransf-T2T 13/3), the central claim that context-aware systems did not outperform the others is not established for this component. The authors should validate the rule with independent native-speaker acceptability judgments on the contested cases, or reframe the word-order results as a descriptive measure of divergence from the FGD preference rather than as translation adequacy.","section":"6.1.1, 7.4, and 8"},{"comment":"The quantitative comparisons are based on very small counts, with focus-proper 'yes/no' observations per system ranging from 0 to 5 and placement observations ranging from 3 to 17; the star charts in Tables 1 and 2 do not report raw frequencies. At these sample sizes, the finding that 'the systems performed with only a minor differences' is indistinguishable from a lack of statistical power, and the absence of a significant difference is not evidence of equivalence. The authors should report exact counts for every cell, and either provide an appropriate statistical test (e.g., Fisher's exact test) or explicitly label the comparison as descriptive. This is load-bearing for the abstract's claim that context-aware systems did not outperform the others.","section":"7.2–7.4 and Tables 1–3"},{"comment":"The word-order analysis assumes that an English subject noun with an indefinite article is contextually non-bound and therefore should appear after the predicate (or as focus proper) in Czech. This assumption is presented as given ('It is assumed'), but the annotators themselves only confirmed the contextual non-boundness for a subset of the automatically selected sentences (85 Yes, 10 No). Given that the whole word-order component depends on this proxy, the paper should report how many of the automatically selected sentences were judged by the annotators as non-bound, and how many of the 'no' placement labels co-occur with a 'no' from the annotators' own contextual-boundness judgment. Without this, the placement results conflate an invalid source-side assumption with a genuine target-side error.","section":"6.1.1 and 7.4"}],"minor_comments":[{"comment":"The inter-annotator agreement is reported as a percentage range and average, but the paper does not state whether this is raw percentage agreement or a chance-corrected measure such as Cohen's kappa; specifying the measure would make the agreement figure interpretable.","section":"7.1"},{"comment":"The phrase 'There were 23 queries in average' should be corrected to 'on average', and the star charts in Tables 1 and 2 would be easier to interpret if exact counts were provided in parentheses next to each star rating.","section":"7.2"},{"comment":"The paper cites Hajičová et al. (1998) for definitions of topic-focus articulation, but it does not justify the specific heuristic that an indefinite article on a subject noun marks contextual non-boundness; a brief justification or pointer to the relevant passage would strengthen the methodological basis.","section":"6.1.1"},{"comment":"The sentence 'This can be attributed to the fact that the systems perform good enough on this task already' is informal; 'perform well enough' would be more appropriate in a journal-style report.","section":"8"},{"comment":"The discussion of word order examples would be clearer if the problematic word order in each Czech example were explicitly highlighted (e.g., with underlining or a gloss), as the current text requires the reader to infer the intended issue.","section":"5.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is best viewed as a resource paper plus an exploratory evaluation. The negative claim about context-aware systems is likely to be cited, so it is important that the word-order metric be validated or the claim carefully qualified. The paper evaluates systems from the authors' own institution, which is not a concern per se, but the framing of the conclusion should not overstate the strength of the evidence given the small counts and the unvalidated rule. The journal should also consider whether the arXiv preprint version has been sufficiently extended for archival publication, since the present version reads like a WMT19 system report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper builds a usable English–Czech test suite for document-level MT evaluation and runs a careful manual evaluation of five WMT19 systems plus the PCEDT reference. That part is worth engaging. But the central negative finding—context-aware systems did not outperform sentence-level systems—rests on a word-order scoring rule that the paper itself applies inconsistently to the human reference, plus very small counts. As it stands, the claim is not established.\n\nThe concrete contribution is the reuse of 101 PCEDT documents with PDTB 3.0 discourse annotations for the WMT19 shared task, and the manual evaluation of AltLexes, connectives, and topic-focus placement. The annotation procedure is described in enough detail to replicate, the interface is shown, and pairwise inter-annotator agreement is reported (80% average). The authors are also candid about limitations. That is real work and a useful resource for English–Czech MT evaluation.\n\nSoft spots, in order. Most important is the word-order metric in Section 6.1.1. The rule says a contextually non-bound English subject introduced by an indefinite article should appear post-verbally, ideally as focus proper, in Czech. The stress-test note is right: in the focus-proper task the human reference is marked “no” in 6 of 7 cases, and in placement it is marked “no” in 3 of 22, while Section 8 says this does not mean the reference is incorrect. A scoring rule that so often flags the human reference is measuring a preferred information-structure ordering, not translation adequacy. This matters because word order is the component where the context-aware systems look worst, so the no-gain conclusion is not safe. Second, the counts are tiny—often 0 to 5 per cell—and no significance testing is done, so “only minor differences” is more a description of the annotation than a measured finding. Third, the star tables in this version render only partially, and no direct link to the test suite is given, which weakens reuse.\n\nCitation patterns are fine. Self-citations to the authors’ own Czech connective work are appropriate, and evaluating their own CUNI systems is disclosed and balanced by online-B and the reference.\n\nWho is this for? People working on MT evaluation or Czech–English discourse. It deserves a serious referee, not because the conclusion is clearly right but because the resource and annotation are worth having. I’d send it out and ask for a revised version that releases the data, reports uncertainty, and either validates the word-order rule against acceptability judgments or softens the comparative conclusion to “no clear difference.”","headline":"The test suite is a useful resource, but the headline no-difference claim rests on a word-order metric that flags the human reference too often to carry it.","tokens_in":10921,"tokens_out":3133,"would_cite":true,"duration_ms":32214,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Giving neural machine translation systems multiple sentences of input did not yield measurable gains on three discourse-level phenomena in English-to-Czech translation at the 2019 WMT shared task.","keywords":["document-level machine translation","test suite","manual evaluation","discourse connectives","AltLex","topic-focus articulation","word order","English to Czech translation"],"falsifier":"Re-annotate the same outputs with direct human judgments of contextual boundness and preferred word order for each subject, instead of the indefinite-article proxy; if the two measurements disagree substantially, the paper's word-order comparison across systems does not reflect actual information-structure handling.","tokens_in":9964,"feed_emoji":"🧪","tokens_out":9075,"duration_ms":79837,"temperature":0.7,"pith_summary":"The paper presents a manually annotated test suite for evaluating document-level translation quality of English-to-Czech neural machine translation (NMT) systems in the 2019 news translation shared task, focusing on three coherence-related phenomena: topic-focus articulation (subject word order), discourse connectives, and alternative lexicalizations of connectives (AltLexes). Evaluating five systems, including two context-aware models that received multiple sentences of input during decoding, the authors found that the context-aware systems did not outperform the sentence-by-sentence systems on these phenomena. Errors were most frequent in word order related to information structure, and the human reference translation was sometimes scored as worse than the automatic outputs, illustrating the difficulty of literal reference comparison. The suite is offered as a reusable resource for future document-level evaluation.","feed_headline":"Extra context did not improve discourse translation at WMT19","feed_subtitle":"Five systems, two with extra context, end up equal on discourse-level Czech translation.","key_machinery":"The evaluation rests on a test suite of 101 documents from a parallel English-Czech treebank, with discourse-relation annotations taken from a discourse treebank. For connectives and AltLexes, trained annotators used a questionnaire to mark each occurrence as adequate and correctly placed, adequate but misplaced, harmlessly omitted, harmful omitted, or inadequate. For word order, the machinery is a proxy: English sentences whose subject is a noun with an indefinite article are automatically preselected, on the theory that such a subject is contextually non-bound and should therefore be placed after the predicate (as focus proper) in Czech. The annotators then judged whether the Czech translation preserved the non-bound status and placed it appropriately. This proxy carries the argument for information-structure sensitivity.","core_discovery":"The central discovery is that, contrary to the authors' expectation, feeding a Transformer-based NMT system several sentences of context does not measurably improve its handling of followed document-level phenomena in English-to-Czech translation. All five evaluated systems translated AltLexes and discourse connectives at an adequate level in the vast majority of cases, with only minor differences among systems. The clearest shortcomings appeared in word order: systems tended to preserve the English order and did not adapt to Czech information structure, so contextually non-bound subjects were often not placed after the predicate as Czech requires. Because the context-aware systems did not outperform the others even here, the authors conclude that these systems already perform 'good enough' on the task, and that the remaining errors are idiosyncratic and difficult to predict.","pith_inferences":["If the result generalizes, document-level NMT for English-to-Czech should shift from feeding adjacent sentences to modeling explicit discourse structure or longer-range coreference, since a few sentences of context gave no benefit.","The indefinite-article proxy could be directly validated against independently annotated topic-focus status in a larger corpus; such a study would establish whether the word-order comparison is an artifact of the proxy.","The absence of a context benefit may be specific to this language pair and system generation; applying the same suite to a language pair with freer word order could reveal larger context effects.","A scalable automatic metric for discourse-connective adequacy, tested against these manual annotations, would allow document-level evaluation to be automated, though the present study only provides the manual gold standard."],"forward_implications":["Context-aware decoding as currently implemented does not automatically improve discourse-level adequacy for English-to-Czech; AltLexes and connectives are already translated at near-reference quality by all observed systems.","The main differentiator of quality lies in word order and information-structure handling, where all systems—including context-aware ones—show comparable error rates.","Bilingual manual evaluation can rate a human reference translation below automatic output, because a looser, more natural translation can be judged as missing or misplacing the traced expressions.","Because errors appear individually and non-systematically, simple deterministic post-editing rules are unlikely to eliminate them.","The released test suite with discourse annotations gives future shared tasks a reusable instrument for measuring document-level phenomena."],"supporting_citations":[{"why":"supplies the parallel English-Czech documents and the human reference translations used as the test suite.","marker":"Hajič et al. (2012)"},{"why":"provides the discourse annotation manual from which connectives and AltLexes are taken.","marker":"Webber et al. (2019)"},{"why":"defines alternative lexicalizations (AltLexes), the phenomenon one annotation task targets.","marker":"Prasad et al. (2010)"},{"why":"supplies the topic-focus articulation theory underlying the word-order proxy.","marker":"Sgall et al. (1986)"},{"why":"describes the base Transformer architecture that the evaluated systems either reuse or extend to document-level translation.","marker":"Popel (2018)"},{"why":"documents the prior English-to-Czech quality level and motivates the bilingual evaluation design.","marker":"Bojar et al. (2018)"},{"why":"defines contextual boundness and topic-focus articulation terms used in the word-order analysis.","marker":"Hajičová et al. (1998)"}],"fun_headline_variants":["Context fails to improve Czech discourse translation","WMT19 test suite finds no discourse gain from extra context","Five systems, two with context: no discourse translation edge","Document-level NMT: context doesn't fix word order mistakes","Test suite: context-aware NMT no better on discourse phenomena"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The word-order results rest on the assumption that an English subject with an indefinite article is always contextually non-bound and should therefore follow the predicate in Czech; this proxy was validated on only a subset of cases and its underlying theory was not independently tested.","fun_headline_variants_meta":{"raw":{"variants":["Context fails to improve Czech discourse translation","WMT19 test suite finds no discourse gain from extra context","Five systems, two with context: no discourse translation edge","Document-level NMT: context doesn't fix word order mistakes","Test suite: context-aware NMT no better on discourse phenomena"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000722,"raw_usage":{"total_tokens":3139,"prompt_tokens":747,"completion_tokens":2392,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":363,"completion_tokens_details":{"reasoning_tokens":2312}},"tokens_in":363,"tokens_out":2392,"duration_ms":18225,"temperature":1.0,"reasoning_tokens":2312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:24:46.569239+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the same outputs with direct human judgments of contextual boundness and preferred word order for each subject, instead of the indefinite-article proxy; if the two measurements disagree substantially, the paper's word-order comparison across systems does not reflect actual information-structure handling.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the discourse annotation manual from which connectives and AltLexes are taken."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines alternative lexicalizations (AltLexes), the phenomenon one annotation task targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the topic-focus articulation theory underlying the word-order proxy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"describes the base Transformer architecture that the evaluated systems either reuse or extend to document-level translation."}],"review_version":1}