{"id":"47596ca2-d53a-49ae-8c0b-a3cf3013f508","arxiv_id":"2508.10421","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 900 annotated translation pairs from nine MT systems, the best system still mistranslates Chinese idioms in 28% of cases, and standard metrics miss these errors (Pearson correlation below 0.48).","lead":"This paper introduces IdiomEval, a benchmark with a detailed error taxonomy for Chinese idiom translation, and shows that leading systems such as GPT-4 still mistranslate idioms in 28% of tested cases. It also finds that standard machine-translation metrics agree poorly with human judgments of idiom quality, and builds detectors that catch such errors.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unreported inter-annotator agreement and idiom sampling make the 28% error rate and sub-0.48 metric correlations unsupported.","rationale":"The central claim is a measurement claim: the 28% error rate, the metric correlations, and the F1=0.68 detector are all outputs of a human-annotation process. The weakest link is the validity of that process. The abstract provides no inter-annotator agreement, no sampling method, and no definition of correct for idioms with multiple acceptable translations; this is exactly the reader's weakest_assumption. Since the supplied full text is unreadable, I cannot determine whether the body resolves this, but the abstract alone does not. A noisy reference standard would also attenuate the reported Pearson correlations, making the claim 'metrics measure idiom quality poorly' potentially overstated. The paper is not internally inconsistent as far as can be checked, and the abstract's claims are coherent, but the measurement foundation is not yet established. The reader's UNVERDICTED verdict is the honest outcome; my concern reinforces it rather than changing it, so the verdict should remain UNCHANGED.","tokens_in":802,"tokens_out":711,"duration_ms":51793,"concrete_test":"1) Draw a stratified random 100-pair subset from IdiomEval (25 per domain) and have two independent bilingual annotators re-label all 100 using the paper's taxonomy, blinded to the original labels; compute Cohen's kappa. If kappa < 0.6, the 28% error rate and metric correlations are not annotation-stable. 2) Recover or request the sampling procedure: how the idiom list was compiled, how source sentences were selected per domain, and whether the 900 pairs are unique sources or repeated across systems. If idioms were handpicked or source sentences came from only a few articles, the 'ubiquitous failure mode' generalization is unsupported. Ideally, rerun the headline comparison on a fresh idiom sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim—GPT-4 makes errors on 28% of Chinese idiom translations and standard metrics cannot detect idiom quality (Pearson <0.48)—rests entirely on the 900 human-annotated pairs in IdiomEval. The abstract specifies the taxonomy (incorrect, literal, partial, missing) and the four domains, but it gives no inter-annotator agreement, no sampling frame for idioms or source texts, and no definition of an 'acceptable' translation for idioms that admit multiple correct renderings. Without these, boundaries between 'partial' and 'literal' are subjective, and the error rate and correlations are not stable measurements. The supplied full text is corrupted mojibake, so the body cannot be checked for whether such details exist. This is not a claim of fraud; it is an unresolved validity threat that is logically upstream of every reported number. If raters disagree on categories, the 28% error rate and the F1=0.68 detector could change materially.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IdiomEval, a framework with an error taxonomy for Chinese idiom translation, and reports a human-annotated dataset of 900 translation pairs produced by nine MT systems, including GPT-4o and Google Translate, across four domains (web, news, Wikipedia, social media). The headline empirical claims are that contemporary systems mistranslate Chinese idioms at a high rate (the best system, GPT-4, errs in 28% of cases), standard automatic metrics correlate poorly with human judgments (Pearson < 0.48), and an error-detection model achieves F1 = 0.68. As received, however, the full text is corrupted mojibake, so only the abstract and a few recoverable fragments can be evaluated. The validity of every reported number depends on the annotation methodology, which cannot be inspected in the supplied text.","tokens_in":13995,"tokens_out":2513,"duration_ms":30961,"significance":"If the claims hold, this paper addresses a real and under-studied failure mode in LLM-based MT: idiom translation quality is largely invisible to standard metrics. The proposed resource (900 annotated pairs, nine systems, four domains) and error taxonomy (incorrect, literal, partial, missing) would be useful to the community. The paper also makes a practical claim that idiom errors can be detected automatically at F1 = 0.68. However, the significance is currently conditional: the headline error rates and correlations are only as reliable as the human annotations, and the detector's score is only meaningful if evaluated on held-out data. The manuscript does not provide the evidence needed to assess either point.","major_comments":[{"comment":"The central empirical claims (28% error rate, Pearson < 0.48, F1 = 0.68) are all outputs of the human annotation process described only as 'We annotate 900 translation pairs ... across four domains.' The abstract reports no inter-annotator agreement, no sampling procedure for selecting idioms or source texts, and no definition of what counts as an acceptable translation for idioms that admit multiple correct renderings. Without these, the boundaries between 'literal,' 'partial,' and 'incorrect' are subjective, and the reported rates and correlations are not stable measurements. Please report agreement statistics and the annotation guidelines, including how multiple acceptable translations were handled.","section":"Abstract"},{"comment":"The statement 'we thus develop improved models that achieve F1 scores of 0.68 for detecting idiom translation errors' does not specify whether the detector was evaluated on a held-out set disjoint from the 900 training/annotation instances. If the F1 is computed on the same annotations used to train the detector, it is a self-evaluation and does not support the implied generalization. Please provide the train/test split, the number of instances, and the precision/recall/F1 for each error category.","section":"Abstract / F1 detector"},{"comment":"The supplied full text is corrupted mojibake; essentially none of the methodology, experiments, tables, or equations can be read. Consequently, the annotation procedure, system list, metric computation, and taxonomy definitions cannot be verified. This is load-bearing because the paper's claims are empirical and depend entirely on details that are not visible. A clean, readable manuscript must be provided before the work can be evaluated.","section":"Full text (all sections)"}],"minor_comments":[{"comment":"The abstract mentions 'GPT-4o' among the evaluated systems but then says 'The best-performing system, GPT-4.' Please clarify whether GPT-4 and GPT-4o are distinct entries and list all nine systems in the main text.","section":"Abstract"},{"comment":"The 28% error rate for GPT-4 is reported without a confidence interval or the number of items it is based on; given the centrality of this number, a simple CI would help assess stability.","section":"Abstract"},{"comment":"'Improved models' is plural but only one F1 value is reported; specify the model architecture(s) and whether F1 is macro-averaged or computed on pooled categories.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The chief barrier is the corrupted full text; this may be an artifact of the submission pipeline rather than the authors' intent, but as received the manuscript cannot be scientifically evaluated. If the authors can provide a clean version, the substantive issues (inter-annotator agreement, sampling, held-out evaluation of the detector) are addressable within a normal revision. There is no apparent novelty or scope concern beyond the need for these methodological details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—quick take on arXiv:2508.10421. The genuinely new thing here is the first systematic snapshot of Chinese idiom translation quality across modern LLM-based systems, built on a human-annotated benchmark (IdiomEval) with a four-way error taxonomy. The headline numbers—best system (GPT-4) errors on 28% of cases, standard metrics correlate below 0.48 with human judgments—are plausible and, if they hold, give the field a target to beat. The abstract is coherent and the design is standard: nine systems, four domains, 900 pairs, human labeling. That is a legitimate contribution.\n\nWhat I couldn't do is check the details. The full text I received is corrupted mojibake, so I only have the abstract. That matters because the abstract omits the three pieces I'd need to trust the numbers cold: inter-annotator agreement, the sampling procedure for idioms and texts, and a working definition of an \"acceptable\" translation for idioms that have multiple correct renderings. None of these omissions is damning in an abstract—plenty of good papers leave them for the body—but they are load-bearing for a human-annotation study. If the body reports IAA and shows the taxonomy, this is a solid paper. If it doesn't, the 28% and the sub-0.48 correlations are provisional.\n\nOne more soft spot, also verification-dependent: the improved model that hits F1=0.68 for detecting idiom errors is trained on the authors' own labels. That's fine if there's a held-out test set with human gold; it's circular if there isn't. The abstract doesn't say.\n\nOn the citation pattern, the abstract frames this as a response to a gap (\"little is known\"), which is exactly the right framing for a new benchmark, and I see no red flags from the outside. Self-citation is not an issue here.\n\nBottom line: this is a useful paper for anyone working on MT evaluation or Chinese NLP, and it deserves a proper referee rather than a desk rejection. The referee's job is to pin down annotation reliability and the detector's held-out evaluation, both of which should be checkable in the body. If those check out, it's an incremental but real benchmark with a memorable negative result. If they don't, the numbers are a headless measurement. I'd send it to review and ask those two questions; my guess is the body answers them.","headline":"Useful first benchmark for Chinese idiom translation, but every headline number is provisional until the body confirms annotation reliability and a held-out detector.","tokens_in":14468,"tokens_out":2466,"would_cite":true,"duration_ms":25164,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The best translation systems tested, including GPT-4, mistranslate Chinese idioms in 28% of cases, and standard automatic metrics cannot detect these failures.","keywords":["Chinese idioms","machine translation","large language models","translation evaluation","error taxonomy","human annotation","figurative language","IdiomEval"],"falsifier":"Ask a second set of annotators to re-label the same 900 translation pairs with the same four-category taxonomy and compute inter-annotator agreement; if agreement is low, or if the 28% error rate for GPT-4 moves materially, the headline results are annotation-dependent rather than a stable property of the systems.","tokens_in":13681,"feed_emoji":"🌐","tokens_out":9310,"duration_ms":87894,"temperature":0.7,"pith_summary":"The paper sets out to show that Chinese idiom translation is a distinct, measurable failure mode of machine translation that current quality checks do not capture. It builds IdiomEval, a set of 900 human-annotated translation pairs from nine systems (including GPT-4 and Google Translate) across web, news, Wikipedia, and social media, and labels each output with one of four error types: incorrect, literal, partial, or missing. The headline result is that the best system tested, GPT-4, still gets 28% of these idioms wrong, while standard automatic metrics correlate with human idiom ratings at Pearson below 0.48—so a human-looking score can hide a mistranslated idiom. A reader should care because, if this holds, idiom quality is an invisible but common gap in every LLM-based translation pipeline, and fixing it requires a dedicated evaluation instrument rather than better generic metrics.","feed_headline":"28% of Chinese idioms are mistranslated by the best translation model","feed_subtitle":"A 900-pair human audit shows standard metrics miss idiom errors; a new detector flags them at F1 0.68.","key_machinery":"IdiomEval is the central instrument: 900 human-annotated translation pairs from nine systems across four domains, organized by a four-way error taxonomy (incorrect, literal, partial, missing) that turns an anecdotal sense that 'idioms are translated badly' into countable categories. The taxonomy does the work: it produces the 28% error rate, provides the human ratings against which standard metrics correlate below 0.48, and supplies labeled training data for the $F_1=0.68$ error detector.","core_discovery":"On the paper's own terms, the discovery is that Chinese idioms expose a blind spot in modern translation systems. Across 900 hand-annotated translation pairs produced by nine systems, every system examined makes errors that fall into four categories—incorrect, literal, partial, and missing—and the best-performing system, GPT-4, errs in 28% of cases. Existing automatic evaluation metrics track human idiom-quality judgments poorly, with Pearson correlation below 0.48, so these failures would not show up in routine evaluations. The paper additionally shows that a detector trained on the annotation data reaches $F_1=0.68$, indicating that automatic error detection is possible but far from perfec","pith_inferences":["The paper samples four written domains, so the 28% rate is not shown to extend to spoken, literary, or domain-specialized Chinese; if it did, idiom errors would be an even more widespread everyday problem than the paper demonstrates.","A natural extension the paper leaves implicit is to test whether the same blind spot appears for idioms in other languages, which would turn a Chinese-specific result into a general property of how large language models handle figurative language.","One testable consequence: a translation system prompted or trained to check for the four error types before output should reduce the 28% error rate; that experiment is not in the paper.","The low metric correlations imply idiom quality should become its own reporting dimension in translation evaluations, rather than being folded into a single overall score."],"forward_implications":["If the 28% figure holds, GPT-4 and comparable systems silently mistranslate more than one in four Chinese idioms in ordinary text.","Because standard metrics correlate below 0.48 with human idiom ratings, systems optimized on those metrics can appear high-quality while regressing on idioms.","The $F_1=0.68$ detector offers a concrete way to flag suspect idiom translations for human review even when overall translation scores look good.","The four error types imply different fixes: literal translations need meaning recovery, partial translations need completeness, missing translations need detection, and incorrect translations need replacement."],"supporting_citations":[],"fun_headline_variants":["GPT-4 mistranslates 28% of Chinese idioms","Chinese idioms expose translation blind spot","Standard metrics fail on Chinese idiom quality","Detector catches Chinese idiom errors, F1 0.68","Even GPT-4 stumbles on 28% of Chinese idioms"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The measurement stands or falls on the human annotations: the abstract reports no inter-annotator agreement, no sampling procedure for the idioms or domains, and no definition of 'correct' for idioms with several acceptable renderings, so if raters disagree, the 28% error rate, the correlation figures, and the detector's $F_1$ all shift.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 mistranslates 28% of Chinese idioms","Chinese idioms expose translation blind spot","Standard metrics fail on Chinese idiom quality","Detector catches Chinese idiom errors, F1 0.68","Even GPT-4 stumbles on 28% of Chinese idioms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1314,"prompt_tokens":700,"completion_tokens":614,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":547}},"tokens_in":444,"tokens_out":614,"duration_ms":6668,"temperature":1.0,"reasoning_tokens":547,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:26:16.819679+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask a second set of annotators to re-label the same 900 translation pairs with the same four-category taxonomy and compute inter-annotator agreement; if agreement is low, or if the 28% error rate for GPT-4 moves materially, the headline results are annotation-dependent rather than a stable property of the systems.","supporting_citations":[],"review_version":1}