{"id":"a87b03a1-49dd-4a8f-a48b-aa6b54a3211b","arxiv_id":"2412.11732","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The WMT 2024 literary translation shared task finds that domain-enhanced systems lead in d-BLEU for Chinese-English, but human evaluators rank NLP2CT-UM and SJTU-LoveFiction at the top.","lead":"This paper presents the results of the second WMT shared task on discourse-level literary translation, covering Chinese to English, German, and Russian. It finds that most participating systems beat strong baselines on document-level BLEU for Chinese-English, while human and automatic rankings sometimes disagree.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract finding (3) is unsupported: no Constrained Track system appears anywhere in the paper, yet the abstract credits one with near-best performance.","rationale":"I read the paper as a shared-task findings report whose central value is the empirical summary in the Abstract. The reader's weakest assumption concerns the reliability of the human evaluation: only one language direction, a small annotator pool, a placeholder for the number of sentences, and inconsistent evaluator counts. Those are real problems and they support a conditional verdict, but they concern the strength of the ranking evidence rather than a direct contradiction with the data. My independent pass found a more concrete load-bearing issue: the Abstract's third finding asserts a comparison involving a Constrained Track system, yet every system in the paper is flagged as unconstrained. That claim cannot be checked against the reported results, so one of the three headline findings is unsupported. I agree with the reader that the paper is otherwise plausible and that the mechanical evaluation issues warrant correction, but I would make the missing constrained system the primary condition for acceptance. The test I propose is a simple check of the official submission records, which would either confirm that the claim is an error or force the authors to include the missing constrained system and its results. I do not see grounds for rejection because the first two findings appear supported by the tables, and the paper releases data, outputs, and a leaderboard, which gives the community a way to verify the reported numbers. However, the conditional verdict should explicitly require the constrained-track claim to be corrected or evidenced before the paper is accepted.","tokens_in":9308,"tokens_out":2961,"duration_ms":28617,"concrete_test":"Query the official WMT24 literary translation leaderboard and submission log at the URL in the paper for any system with a Constrained Track flag. If no constrained submission exists, delete or revise Abstract finding (3) so that the paper only reports unconstrained systems. If a constrained submission does exist, add it to Table 2 with flag 'J' and include its automatic and human scores in Tables 3–5, then re-evaluate whether its performance is close to the best unconstrained system. In either case, add a note identifying which system, if any, is the 'certain system' behind finding (2), and report the number of human-judged sentences, replacing 'xxx' in §3.2.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's third headline finding, stated in the Abstract, is that 'the performance of the only Constrained Track system was close to that of the best system in the Unconstrained Track.' The body of the paper does not contain any constrained system. Table 2 lists all five submissions and marks every one with flag 'N' (unconstrained); the flag legend even includes 'J' but no row uses it. Section 1 says 'most teams choose Unconstrained Track,' which could imply at least one constrained team, but no constrained system is described in Section 5, no constrained row appears in Tables 3–5, and every system in those tables is marked ⋆ = unconstrained. Thus the comparison that finding (3) reports cannot be reconstructed from the paper's data; it is either an editing leftover or refers to an unlisted system. Because this is one of the three central claims that the Abstract promises, the findings as written overstate what was demonstrated. The first two findings are largely supported by Tables 3 and 5, although claim (2) about a 'significant gap' is not accompanied by any significance test or confidence interval, so its strength is also overstated. The human-evaluation concerns identified by the reader (placeholder 'xxx' in §3.2, inconsistent evaluator count, single language direction) further limit the force of the ranking claims, but the missing constrained system is the more direct problem: it is a stated result with no corresponding evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents the findings of the WMT 2024 shared task on Discourse-Level Literary Translation, covering three language directions (Chinese→English, Chinese→German, Chinese→Russian). The organizers describe the second edition of the GuoFeng Webnovel Corpus, the evaluation protocol (automatic metrics plus human evaluation), the five participating teams, and their systems. The abstract states three headline findings: (1) most participating systems outperform baselines (Llama-MT, Google Translate, GPT-4) in d-BLEU; (2) for a certain system, there is a significant gap between conclusions drawn from human and automatic evaluations; and (3) the only Constrained Track system performed close to the best Unconstrained Track system in both automatic and human evaluation.","tokens_in":9559,"tokens_out":5963,"duration_ms":47420,"significance":"Findings papers for WMT shared tasks serve as community reference points for state of the art in a given translation direction. This paper contributes a new corpus extension (Zh-De, Zh-Ru), baseline comparisons, and system summaries, and it makes the data, outputs, and leaderboard publicly available. If the reported claims were fully supported, the paper would be a useful record for researchers working on discourse-level and literary MT. However, several of the central claims are currently not backed by the presented evidence, and the human evaluation protocol is described with unresolved inconsistencies. The paper's value depends on correcting these issues.","major_comments":[{"comment":"Abstract finding (3) claims that the only Constrained Track system performed close to the best Unconstrained Track system, but no constrained system appears anywhere in the paper. Table 2 lists all five participating teams with the flag 'N' (unconstrained); no row uses the flag 'J', and every system in Tables 3 and 4 is marked with the '⋆' symbol defined as unconstrained. Section 5 describes each team's methods without mentioning any constrained submission. Because one of the three headline findings cannot be reconstructed from the reported evidence, the abstract overstates the paper's demonstrated results.","section":"Abstract and Section 5 / Table 2"},{"comment":"The human evaluation protocol is incompletely specified and internally inconsistent. The text contains the literal placeholder 'For each chapter, we assessed xxx sentences', leaving the evaluation unit undefined. Immediately afterward, the paper says 'We employed two professional evaluators' and then 'We employed three professional evaluators', while Table 1 lists three evaluators (A, B, C). Since the official ranking is said to be based on these human judgments, the number of judgments per system, the evaluation unit, and the consistent evaluator count are load-bearing details that must be reported accurately.","section":"Section 3.2, Human Evaluation"},{"comment":"Abstract finding (2) states that there is a 'significant gap' between human and automatic evaluation conclusions for a certain system, but no such system is identified and no statistical evidence is provided. Section 6.2 reports human scores without any significance test, confidence interval, or explicit comparison between the human and automatic rankings. The word 'significant' implies quantitative support that is absent from the paper, so this headline claim is currently unsubstantiated.","section":"Abstract and Section 6.2"},{"comment":"The ranking information in Table 5 is ambiguous and contradicts the text. The Rank column shows '1 / 2' for NLP2CT-UM and '2 / 1' for SJTU-LoveFiction, without specifying which rank corresponds to General quality and which to Discourse awareness. Section 6.2 states that NLP2CT-UM 'achieved the highest scores in General (3.96) and Discourse (4.00)', but SJTU-LoveFiction has a Discourse score of 4.13, which is higher than 4.00. This inconsistency affects the official human-based ranking, which is the paper's central product.","section":"Table 5 and Section 6.2"},{"comment":"The claim in Section 6.1 that 'the primary systems outperform the baselines' in d-BLEU is not supported by the reported data. In Table 3, NTU has a d-BLEU of 34.6, well below Google's 47.3, and in Table 4 both NLP2CT-UM and SJTU-LoveFiction score below the Google baseline for Zh-De and Zh-Ru. Abstract finding (1) says 'most' systems outperform baselines, which may be salvageable, but the unqualified wording in Section 6.1 is contradicted by the tables and should be corrected to reflect the actual comparisons.","section":"Section 6.1, Tables 3 and 4"}],"minor_comments":[{"comment":"The text says sacreBLEU, chrF, TER, and d-BLEU are calculated 'with two references', but the Zh-De and Zh-Ru test sets are described in Section 2.2 as having only one reference; the metric configuration for each language direction should be clarified.","section":"Section 3.1"},{"comment":"The Zh-Ru row shows 'Test 122 13 15.6K'; the book count is presumably 12 rather than 122, which appears to be a typo.","section":"Figure 1"},{"comment":"Both sections contain 'introduced 1 translation systems'; the noun should be singular.","section":"Sections 5.1 and 5.4"},{"comment":"The in-text citation 'Chinese-Llama-2 7B (Zefeng Du, 2023)' does not match any reference entry; the closest entry is 'Longyue Wang Zefeng Du, Minghao Wu. 2023. Chinese-llama-2.' This should be reconciled in the reference list.","section":"Section 2.3 and References"},{"comment":"The flag legend includes 'J' for constrained systems, but no row uses this flag; the paper should state explicitly that no constrained systems participated in the final evaluation, or if one did, it must be added to the table and described.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to have been prepared hastily: the 'xxx' placeholder in Section 3.2, the two/three evaluator inconsistency, and the unsupported Constrained Track claim in the abstract are all issues that a careful revision must address. If the constrained system genuinely participated, it must be described and included in all tables; if not, the abstract claim must be removed. The ambiguous ranking in Table 5 is especially important because the official ranking is the paper's main deliverable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the WMT 2024 discourse-level literary translation findings paper. What's genuinely new: two new language pairs (Zh-De, Zh-Ru), the GuoFeng corpus v2 release, and the empirical result that several submitted systems beat Google/GPT-4 baselines on d-BLEU for Zh-En. The human eval for Zh-En is also a step beyond last year, and the system descriptions are useful pointers to the participants' own papers. As a benchmark report, it does its job.\n\nThe soft spots are real but concentrated. The abstract's third finding — \"the performance of the only Constrained Track system was close to that of the best system in the Unconstrained Track\" — has no supporting evidence anywhere in the body. Every system in Table 2 is flagged unconstrained (N); no row uses the constrained flag (J). So the claim is either an editing leftover or refers to a system that isn't described. That is a stated result with no corresponding data, and it should be fixed before publication.\n\nSecond, the human evaluation section has mechanical issues: Section 3.2 says \"we assessed xxx sentences\" with the placeholder, and the evaluator count is reported as two in one paragraph and three in the next. Since the official ranking is based on that human eval, these details matter. Also, the Zh-De and Zh-Ru results are only d-BLEU with two teams per direction, so those rankings are thin.\n\nThird, the \"significant gap\" between human and automatic conclusions (finding 2) is not backed by any significance test or confidence interval. It's a qualitative observation. The tables do show a gap — NLP2CT-UM is first on d-BLEU and tied first on human, while SJTU-LoveFiction is second on d-BLEU but first on human discourse — but calling it significant is going beyond the evidence.\n\nMy overall read: the core empirical content (tables 3-5, data release) is plausible and consistent. The abstract overstates what was demonstrated, and the missing constrained system is a clear error. It's not a methodological train wreck; it's a findings report with sloppy editing and one unsupported headline.\n\nWho is this for? People working on literary MT or document-level translation will want the data and the leaderboard. The paper itself is a useful pointer, not a deep analysis. I'd send it to review — shared task findings are standard venue material — but condition on fixing the abstract claim, the placeholder, and the evaluator count. A serious referee would catch these in ten minutes.\n\nRecommendation: engage, but require revisions.","headline":"Useful shared-task report with a real data-release contribution, but the abstract claims a constrained-track result that the paper never shows.","tokens_in":10135,"tokens_out":2279,"would_cite":false,"duration_ms":19670,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Literary-domain tuned systems beat general-purpose baselines on document-level BLEU in WMT 2024, but human judgments disagreed with the automatic ranking for one system.","keywords":["literary translation","discourse-level machine translation","document-level BLEU","human evaluation","GuoFeng Webnovel Corpus","WMT shared task","large language models","Chinese-to-English translation"],"falsifier":"Using the released Test_final outputs and references, recompute d-BLEU per book and compare it with new independent human scores on the same passages: if the human/automatic disagreement disappears or the top ranking flips under a different set of annotators, the paper's central findings would not replicate.","tokens_in":9120,"feed_emoji":"📚","tokens_out":10558,"duration_ms":91466,"temperature":0.7,"pith_summary":"This paper reports the results of the second WMT shared task on discourse-level literary translation, covering Chinese-to-English, Chinese-to-German, and Chinese-to-Russian. Its central claim is empirical: after literary-domain enhancements, most participating systems outperform the Baseline systems (Llama-MT, Google Translate, and GPT-4) in terms of document-level BLEU (d-BLEU); for one system, human and automatic evaluations disagree sharply; and the only Constrained Track system performed close to the best Unconstrained Track system. The paper also releases the GuoFeng Webnovel Corpus V2, along with system outputs and a leaderboard, so the claims can be checked and extended. The result matters because it provides a public, reproducible benchmark for whether discourse-aware literary machine translation is actually improving rather than merely producing superficially fluent sentences.","feed_headline":"Tuned systems beat GPT-4 at webnovel translation","feed_subtitle":"Official ranking rests on human judgment, and one system's standing flipped between human and automatic scores.","key_machinery":"The central mechanism is the evaluation design. The document-level automatic score, d-BLEU, concatenates every sentence in a document into a single line and computes n-gram matches with sacreBLEU against two references. The human rubric gives each window of neighboring sentences two scores from 0 to 5, one for general quality (fluency and adequacy) and one for discourse awareness (consistency, word choice, anaphora, tone); the official ranking is built from these human scores rather than from d-BLEU. The newly added language directions, Chinese-to-German and Chinese-to-Russian, have no sentence-level alignment and only one reference, and because only two teams participated in each direction, those directions were evaluated automatically only.","core_discovery":"The paper's findings, stated in the abstract and supported by the shared-task results, are threefold. First, most participating teams' models, after enhancements aimed at the literary domain, outperform the Baseline systems (Llama-MT, Google Translate, and GPT-4) in terms of d-BLEU, the document-level automated metric computed by concatenating each document's sentences into one line before applying sacreBLEU. Second, for a particular system the conclusions drawn from human evaluation diverge significantly from those drawn from automatic metrics. Third, the only Constrained Track submission was close to the best Unconstrained Track submission under both types of evaluation. The official ranking is based on the overall human judgments, which were collected only for the Chinese-to-English direction using a rubric that scores general quality and discourse awareness separately.","pith_inferences":["The human/automatic disagreement suggests that d-BLEU alone would have produced a different leaderboard; a direct test is to recompute the ranking using only automatic metrics and compare it with the official human-based ranking.","Because the human evaluation covers only one language direction and a small number of annotators, the claimed gap between human and automatic conclusions should be read as a hypothesis about literary MT evaluation rather than a universal property; re-running the rubric on more languages and more annotators would show whether the gap is systematic.","The new Chinese-to-German and Chinese-to-Russian references were created by machine translation with human post-editing, so systems that stay close to Google's n-gram style may be favored by d-BLEU; a human evaluation on those directions could reveal whether Google's top d-BLEU score reflects real perceived quality.","The released system outputs would allow a per-book breakdown to test whether the d-BLEU advantage of primary systems over baselines holds outside books represented in the training data."],"forward_implications":["Literary-domain adaptation on web-novel data yields measurable d-BLEU gains over general-purpose translation baselines in Chinese-to-English.","Document-level and human evaluation are not interchangeable: the same system can appear strong under one and weak under the other.","A constrained system can be competitive with unconstrained systems in this task, so access to external data or models is not a decisive advantage.","For language pairs without sentence-level alignment, document-level automatic scoring is the only evaluation used, so conclusions about Chinese-to-German and Chinese-to-Russian rest on d-BLEU alone, where the Google baseline ranked first."],"supporting_citations":[{"why":"Previous edition of the shared task; defines the GuoFeng Webnovel Corpus V1 and the evaluation setup that this year's task extends.","marker":"Wang et al. 2023c"},{"why":"Defines sacreBLEU, the implementation used for sentence-level BLEU and for the document-level d-BLEU scores with two references.","marker":"Post, 2018"},{"why":"Provides the COMET neural metric used alongside chrF, TER, and BLEU in automatic evaluation.","marker":"Rei et al., 2020"},{"why":"Defines chrF, the character n-gram F-score reported as a sentence-level automatic metric.","marker":"Popovi´c, 2015"},{"why":"Defines TER, the translation-edit-rate metric reported as a sentence-level automatic metric.","marker":"Snover et al., 2006"},{"why":"Source of the Cohen's kappa statistic used to claim annotator consistency (0.86) for the human evaluation.","marker":"McHugh, 2012"}],"fun_headline_variants":["Tuned systems beat commercial baselines in WMT literary task","Human and automatic scores diverge in literary translation ranking","Constrained track entry rivals unconstrained in literary translation","WMT literary translation: top systems outperform GPT-4 and Google"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the human scores, collected by three annotators for the Chinese-to-English direction on a custom 0–5 rubric, are a reliable measure of translation quality; if those scores are noisy or unrepresentative, the official ranking and the claimed human-automatic gap lose their foundation.","fun_headline_variants_meta":{"raw":{"variants":["Tuned systems beat commercial baselines in WMT literary task","Human and automatic scores diverge in literary translation ranking","Constrained track entry rivals unconstrained in literary translation","WMT literary translation: top systems outperform GPT-4 and Google"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2310,"prompt_tokens":805,"completion_tokens":1505,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":421,"completion_tokens_details":{"reasoning_tokens":1448}},"tokens_in":421,"tokens_out":1505,"duration_ms":11037,"temperature":1.0,"reasoning_tokens":1448,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:37:18.899605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Using the released Test_final outputs and references, recompute d-BLEU per book and compare it with new independent human scores on the same passages: if the human/automatic disagreement disappears or the top ranking flips under a different set of annotators, the paper's central findings would not replicate.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines sacreBLEU, the implementation used for sentence-level BLEU and for the document-level d-BLEU scores with two references."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the COMET neural metric used alongside chrF, TER, and BLEU in automatic evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines chrF, the character n-gram F-score reported as a sentence-level automatic metric."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines TER, the translation-edit-rate metric reported as a sentence-level automatic metric."}],"review_version":1}