{"id":"7de3a6f9-b49d-47e3-adac-891e8e60f5da","arxiv_id":"2504.14804","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of document-level machine translation evaluation that catalogs existing metrics, describes their weaknesses, and proposes future directions without introducing a new metric or result.","lead":"This paper reviews automatic evaluation metrics for document-level machine translation, from traditional lexical metrics to LLM-based judges. It argues that current metrics remain limited by reference diversity, sentence alignment dependencies, and judge bias.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.3's judge-inaccuracy claim rests on one anecdote that equates token-count disparity with omission; without human adjudication or a reported judge setup, the 8.5 versus 8 score does not demonstrate inaccuracy.","rationale":"The reader's weakest assumption points at the same example, and I agree the methodology is missing; I would sharpen it further. The load-bearing issue is not only that the judge setup is unreported, but that the example never establishes the ground-truth ranking it claims the judge got wrong. Token count is an inadequate proxy for omission, so even the anecdote, as presented, cannot support the 'inaccuracy' label. This concern does not change the reader's CONDITIONAL verdict: the survey's overall taxonomy and the cited bias result give the central claim independent support, but the specific 'inaccuracy with different lengths' finding must be either removed, relabeled as an anecdote, or replaced with systematic evidence. No change is needed relative to the reader's verdict; agreement is partial because the reader emphasized representativeness and protocol, while I emphasize the missing ground truth.","tokens_in":9554,"tokens_out":7597,"duration_ms":69454,"concrete_test":"Run a blind human-adjudication study on Table 1: give the source, reference, V3 output, and R1 output to at least two professional translators in counterbalanced order, with an error typology that separates omissions from legitimate compression, and record which output each annotator prefers and why. If the human rank order does not clearly place V3 above R1, the inaccuracy example collapses; if it does, rerun the LLM judge with a disclosed prompt and model version over 20 position-shuffled trials to check whether the 8.5 versus 8 result is stable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.4 asserts that neither d-Comet nor LLM-as-a-judge can fully align with human evaluation, and the abstract lists 'inaccuracy' as one of the three core LLM-as-a-judge challenges. The only evidence for that inaccuracy claim is the Table 1 anecdote. The example fails to demonstrate judge error for two reasons. First, token-count disparity (R1: 525 tokens versus reference: 879 tokens) is treated as proof of 'obvious omissions,' but a shorter translation can still be faithful; the table does not mark any omission, and the displayed R1 text is mostly a compression of the reference. Second, the true quality ranking is assumed rather than measured: no human adjudication, error annotation, judge-model version, prompt, temperature, number of runs, or position shuffling is reported, even though Section 2.2 recommends shuffling and averaging. The associated sentence that 'longer texts are more prone to omissions' is also internally inconsistent, because the allegedly omission-prone R1 output is the shorter one. Therefore, the paper's claim that LLM judges are inaccurate 'especially when dealing with translation results of different lengths' is not supported by the presented evidence. The survey's other support—the cited self-preference bias study and d-Comet's alignment dependency—is reasonable, so the overall central claim is not overturned; the unsupported example is a scope-of-claim problem.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a short survey of automatic evaluation metrics for document-level machine translation. It organizes the field into reference-based and reference-free schemes, and into traditional (BLEU-style), model-based (COMET, BERTScore, d-COMET), and LLM-based (LLM-as-a-judge) metrics. It then lists four challenges: lack of reference diversity, d-COMET's dependence on sentence-level alignment, LLM-as-a-judge bias/inaccuracy/lack of interpretability, and the gap between automatic metrics and human evaluation. Finally it proposes future research directions such as reducing sentence-level dependency, multi-granular evaluation, and training specialized LRMs for MT evaluation. The paper's most specific empirical contribution is a single example (Table 1) intended to show LLM-as-a-judge inaccuracy.","tokens_in":9804,"tokens_out":5758,"duration_ms":49537,"significance":"If the challenge taxonomy and the empirical claim were properly supported, the paper would be a useful concise reference for practitioners entering document-level MT evaluation. The paper usefully brings together the distinction between sentence- and document-level evaluation, names the key failure modes, and cites the self-preference bias work [25] to substantiate the bias challenge. Its main novelty is the claim, based on the authors' own example, that LLM-as-a-judge is inaccurate for outputs of different lengths; because this claim is not adequately evidenced, the survey's central message is only partially supported. As a survey, it provides structure but not a systematic methodology or comparative synthesis.","major_comments":[{"comment":"The paper's central empirical claim that LLM-as-a-judge is 'highly inaccurate' rests on a single anecdote. The judge setup is not reported (prompt, model version, temperature, number of runs, position shuffling), although Section 2.2 states that multiple runs and shuffling are required. Moreover, the ranking is asserted as 'clearly unreasonable' without human adjudication or error annotation. The table itself shows only token counts and scores; no omissions are marked, so the 8.5 versus 8 score does not demonstrate an evaluation error. To support the claim, provide a controlled evaluation with human adjudication and full judge configuration, or explicitly downgrade the claim to an anecdotal observation.","section":"Section 3.3, Table 1"},{"comment":"The sentence 'using a Large Reasoning Model (LRM) for translation, longer texts are more prone to omissions' is contradicted by the accompanying example: the R1 output that allegedly has 'obvious omissions' is 525 tokens, shorter than the V3 output (802 tokens) and the reference (879 tokens), and no long-text comparison is conducted. Token-count disparity alone cannot distinguish omission from concise paraphrase. Consequently, the concluding sentence that inaccuracy is present 'especially when dealing with translation results of different lengths' is not a supported inference from the presented evidence.","section":"Section 3.3"}],"minor_comments":[{"comment":"The title contains stray spaces in 'T ranslation' and 'T rends'; these should be corrected throughout.","section":"Title"},{"comment":"The abstract lists 'bias, inaccuracy, and lack of interpretability' as challenges, but the conclusion lists only 'bias and lack of interpretability'; align the two lists.","section":"Abstract and Section 5"},{"comment":"The text says 'Comet includes several variants, such as Comet20 and Comet22' but does not specify precisely which models these abbreviations refer to; please expand or add citations.","section":"Section 2.2"},{"comment":"The statement that 'in some machine translation research, the LLM-as-a-judge evaluation method has also been applied' would benefit from concrete citations to that research.","section":"Section 2.2"},{"comment":"The claim that a single reference translation is insufficient for LLM-generated document-level translations is asserted without citations; supporting literature should be provided.","section":"Section 3.1"},{"comment":"Reference [26] ('Text style transfer back-translation') lacks a publication venue and DOI; the entry should be completed.","section":"References"},{"comment":"The claim about the discrepancy between metrics and human evaluation is not supported by citations to WMT Metrics shared-task results or similar studies; adding such references would strengthen the survey.","section":"Section 3.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is quite short for a journal-level survey and does not report a systematic search or selection process for the literature covered. The main technical issue is the unsupported experimental claim in Section 3.3; if the authors can either provide a proper controlled evaluation or reframe the claim as an anecdotal observation, the paper becomes publishable as a position or overview paper. The editors may also want to consider whether the depth of coverage, such as the absence of discussion of recent WMT Metrics tasks and human-evaluation protocols, matches the journal's expectations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper is a survey, not a research contribution. The tripartite taxonomy (traditional/model-based/LLM-based) and the challenges (reference diversity, sentence alignment dependency, judge bias/inaccuracy/interpretability) are all assembled from cited work, and the only new artifact is the Table 1 example. That's fine for an overview, and the prose is clear and accurate in most places. Someone entering document-level MT evaluation would get a reasonable map and a good pointer list.\n\nThe soft spot is exactly where the reader put it: Section 3.3. The claim that LLM-as-a-judge is 'inaccurate' rests on a single example where R1 (525 tokens) beat V3 (802 tokens) against an 879-token reference. That demonstrates nothing by itself. A shorter translation can be faithful; the table doesn't mark concrete omissions; and there is no judge model version, prompt, temperature, number of runs, position shuffling, or human adjudication. The paper's own Section 2.2 tells practitioners to shuffle and average, but the example doesn't show that was done. The surrounding sentence that 'longer texts are more prone to omissions' is also confusing, since the allegedly omission-prone output is the shorter one. So the abstract's 'inaccuracy' as a core challenge is not supported by the paper's evidence. To be fair, the bias claim is backed by a real cited study, and d-Comet's alignment dependency is a real limitation, so the overall message that automatic metrics have problems survives. The fix is easy: label Table 1 as a motivating anecdote or replace it with a small systematic study.\n\nI don't see a novelty problem beyond what the reader flagged. The self-citations are to earlier MT systems, which is not circular. The survey's main weakness is overclaiming from one anecdote. If that's fixed, it's a decent overview. Honestly, I wouldn't cite it in my own work, but I'd be happy to see a revised version in print. It deserves a serious referee: the topic matters, the authors know it, and the one real flaw is fixable.","headline":"Competent survey of document-level MT evaluation whose one empirical claim about LLM-as-a-judge inaccuracy is an under-supported anecdote; useful as an entry point, not a research result.","tokens_in":10379,"tokens_out":3227,"would_cite":false,"duration_ms":26743,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that neither d-Comet nor LLM-as-a-judge can fully align with human evaluation of document-level translation, and that LLM judges show bias, inaccuracy, and opacity.","keywords":["Automatic Evaluation Metrics","Document-level Translation","d-Comet","LLM-as-a-judge","Machine Translation Evaluation","Reference Diversity","Evaluation Bias","Large Language Models"],"falsifier":"A reader could rerun the Table 1 comparison with a fixed prompt, a fixed judge model version, many repetitions with randomized output order, and human adjudication; if the omission-filled 525-token output no longer scores above the 802-token output, the paper's specific evidence for judge inaccuracy collapses.","tokens_in":9343,"feed_emoji":"📊","tokens_out":6034,"duration_ms":49735,"temperature":0.7,"pith_summary":"The paper argues that automatic evaluation of document-level machine translation has not yet caught up with the translation systems it is meant to judge. Its central claim is that neither d-Comet, a context-aware extension of the Comet metric, nor LLM-as-a-judge, where a large language model scores or ranks translations, can fully align with human evaluation. The paper also identifies the reasons: reference translations lack diversity, document-level metrics depend on sentence-level alignment that LLM outputs often break, and LLM judges display self-preference bias, inaccuracy across translation lengths, and scores that are hard to interpret. A fair reader should care because evaluation metrics guide model development and system comparison, so a metric that rewards omission-heavy output can send development in the wrong direction. The paper closes by proposing future directions: less sentence-alignment dependence, multi-level and multi-granular evaluation, and evaluation models trained specifically for judging machine translation.","feed_headline":"LLM-as-a-judge can rank an omission-filled translation higher","feed_subtitle":"A survey finds d-Comet and LLM judges still fall short of human evaluation of document-level translations.","key_machinery":"The two load-bearing mechanisms are d-Comet and LLM-as-a-judge. d-Comet is the method of extending any pretrained sentence-level metric, such as Comet, to the document level by encoding surrounding context; in this paper it stands for the entire class of reference-based document-level metrics and their dependence on crisp sentence alignment. LLM-as-a-judge is the practice of using a large language model to score, rank, or select machine translation outputs; it stands for the reference-free branch of evaluation. The paper works by testing the implicit promises of these two mechanisms against practical constraints: alignment fragility for d-Comet, and bias, inaccuracy, and opacity for LLM-as-a-judge.","core_discovery":"On the paper's own terms, the central claim is a negative result about the current state of the field: no existing automatic metric is a reliable substitute for human evaluation of document-level translation. d-Comet is limited because it requires source, reference, and translation to be split into the same number of sentences and aligned sentence by sentence, while LLM translations often merge or re-segment sentences. LLM-as-a-judge is limited because it is biased toward its own outputs, can be inaccurate when translation lengths differ, and gives little interpretable justification for its scores. The paper supports the inaccuracy claim with a single scored example in Table 1: DeepSeek R1's 525-token translation, which has obvious omissions, received 8.5 from the LLM judge, while DeepSeek V3's 802-token translation, close in length to the 879-token reference, received 8. The paper therefore concludes that neither metric can fully align with human evaluation.","pith_inferences":["If LLM judges systematically prefer shorter outputs, document-level leaderboards may be silently rewarding systems that compress and omit; this could be tested by scoring many outputs with controlled length ratios against human judgments.","The alignment-fragility diagnosis points to a concrete extension: evaluate documents directly on whole-document representations, or use automatic segmentation and alignment, so LLM re-segmentations no longer break the metric.","The paper's error-type proposal suggests a falsifiable design: judges that first identify omissions should penalize omission-laden outputs more than holistic judges do, and that difference could be measured on Table 1-style examples.","Because judge-model versions and prompts are not specified, the reported 8.5-versus-8 gap is probably unstable across judge versions; re-running with several judge models would quantify that instability."],"forward_implications":["System comparisons built on these metrics may misreport quality, since an LLM judge can rank an omission-laden output above a more complete one.","d-Comet is reliable only when documents can be cleanly segmented and aligned sentence by sentence, which LLM-generated translations frequently violate.","Single-reference evaluation will under-reward valid alternative translations, making reported quality worse than actual quality for high-diversity LLM outputs.","Moving to multi-level, multi-granular evaluation with error-type identification, rather than one holistic score, is the paper's proposed route to more interpretable judgments.","Training a specialized evaluation model on human reasoning and scores is the paper's proposed route to more accurate and reliable judgments."],"supporting_citations":[{"why":"It supplies the n-gram overlap baseline that the paper contrasts with semantic and model-based metrics.","marker":"[15]"},{"why":"It defines Comet, the model-based metric that d-Comet extends to the document level.","marker":"[16]"},{"why":"It provides the d-Comet method for converting any pretrained metric into a document-level metric by encoding context.","marker":"[22]"},{"why":"It defines the LLM-as-a-judge evaluation approach used throughout the paper's analysis.","marker":"[7]"},{"why":"It supplies the evidence that LLM judges have self-preference bias, a central charge in Section 3.3.","marker":"[25]"},{"why":"It motivates the Chain-of-Thought-style, multi-granular evaluation proposed as a future direction.","marker":"[10]"},{"why":"It provides BertScore as another model-based metric in the current-state survey.","marker":"[31]"}],"fun_headline_variants":["LLM judges score omission-filled translation higher than complete one","No automatic metric matches human evaluation for document-level translation","d-Comet and LLM judges fall short on document translation evaluation","Omission-filled translation gets higher LLM judge score than accurate one"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inaccuracy claim rests on a single scoring example in Table 1, where a shorter, omission-prone output scored higher than a fuller one; if that example is not representative, the paper's empirical case for LLM judge inaccuracy is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["LLM judges score omission-filled translation higher than complete one","No automatic metric matches human evaluation for document-level translation","d-Comet and LLM judges fall short on document translation evaluation","Omission-filled translation gets higher LLM judge score than accurate one"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000421,"raw_usage":{"total_tokens":2187,"prompt_tokens":992,"completion_tokens":1195,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":1125}},"tokens_in":608,"tokens_out":1195,"duration_ms":8604,"temperature":1.0,"reasoning_tokens":1125,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:39:26.092150+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could rerun the Table 1 comparison with a fixed prompt, a fixed judge model version, many repetitions with randomized output order, and human adjudication; if the omission-filled 525-token output no longer scores above the 802-token output, the paper's specific evidence for judge inaccuracy collapses.","supporting_citations":[{"cited_title":"Embarrassingly Easy Document-Level MT Metrics: How to Convert Any Pretrained Metric Into a Document-Level Metric","cited_arxiv_id":"2209.13654","evidence_quote":"It provides the d-Comet method for converting any pretrained metric into a document-level metric by encoding context."}],"review_version":1}