{"id":"595cb9c6-b8c3-4357-9df3-d34da53718f8","arxiv_id":"2605.15282","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Large-scale analysis of literary translations reveals a consistent negative correlation between fluency (measured via POS n-gram translationese classifier) and faithfulness (COMET-KIWI), controlled for length, across human and machine systems.","lead":"The study analyzes over 130,000 paragraphs from novels and finds a negative correlation between how natural a translation sounds and how well it preserves the original meaning, after accounting for paragraph length. This tradeoff appears in both human and Google Translate outputs but is weaker for one AI system, suggesting that fluency and accuracy may pull in opposite directions when evaluating literary translations.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"POS n-gram translationese classifier may capture syntactic artifacts rather than target fluency, potentially inducing spurious negative correlation with COMET-KIWI","rationale":"The reader's weakest assumption directly identifies the metric-validity issue for literary paragraphs. The proposed test isolates whether the automatic proxies align with human judgments on the same texts, which would confirm or refute whether the negative correlation reflects a real tradeoff or a measurement artifact. This keeps the critique focused on the load-bearing assumption without questioning author intent or requiring full re-analysis of the 130k corpus.","tokens_in":1653,"tokens_out":405,"duration_ms":36106,"concrete_test":"On a held-out set of 500 literary paragraphs, obtain human fluency ratings (1-5 scale for native-likeness and stylistic naturalness) and faithfulness ratings (1-5 for semantic and stylistic fidelity to source). Compute Spearman correlation between (a) human fluency vs. POS n-gram classifier score and (b) human faithfulness vs. COMET-KIWI; if either correlation is below 0.4 or if the partial correlation between the two human ratings (controlling for length) is positive while the automatic metrics show negative, the proxy validity fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the paragraph-level POS n-gram classifier validly proxies target-language fluency (original-likeness) while COMET-KIWI independently measures semantic faithfulness. However, POS n-grams primarily encode syntactic patterns; literary fluency also depends on lexical choice, collocations, and stylistic naturalness not captured by POS sequences. If the classifier and COMET-KIWI both respond to the same underlying syntactic or length-related features (even after linear control for paragraph length), the observed negative correlation could be an artifact of metric construction rather than evidence of a genuine tradeoff. This risk is higher for human and Google Translate outputs than for TranslateGemma if the latter's outputs differ systematically in syntactic distribution.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript examines the relationship between fluency and faithfulness in literary translation using a dataset of 130,486 paragraphs from 106 novels across 16 source languages. Fluency is operationalized as original-likeness via a paragraph-level POS n-gram translationese classifier, while faithfulness is measured with COMET-KIWI. After controlling for paragraph length, the authors report a consistent negative correlation for human and Google Translate outputs that is weaker and frequently non-significant for TranslateGemma. The work concludes that segment length matters for automatic evaluation and that a fluency-faithfulness tradeoff exists in literary translation.","tokens_in":1818,"tokens_out":536,"duration_ms":40932,"significance":"If the central measurements are valid, the study offers empirical support for a tradeoff between target fluency and source faithfulness in literary text, with implications for both human and machine translation evaluation. The large scale of the dataset and the inclusion of multiple translation sources are strengths. The result is most consequential if the POS n-gram classifier can be shown to capture fluency beyond syntactic artifacts.","major_comments":[{"comment":"Methods section (classifier description): The claim that the paragraph-level POS n-gram translationese classifier validly measures target-language fluency (original-likeness) is load-bearing for the negative-correlation result. POS sequences primarily encode syntactic distributions; literary fluency also depends on lexical choice, collocations, and stylistic naturalness. Without validation against human fluency ratings on literary paragraphs or an ablation showing that the classifier is not reducible to length or syntax alone, the observed negative correlation with COMET-KIWI risks being an artifact of shared sensitivity to syntactic or length-related features.","section":"Methods"},{"comment":"Results (correlation tables/figures): The length control is described at a high level, but the manuscript does not report the exact regression specification, variance inflation factors, or residual diagnostics. If residual length effects or genre-specific syntactic patterns remain, they could induce the reported negative correlation independently of any genuine fluency-faithfulness tradeoff.","section":"Results"}],"minor_comments":[{"comment":"Abstract: The sentence on TranslateGemma could explicitly note the sample size or number of languages to allow readers to gauge the power of the non-significant findings.","section":"Abstract"},{"comment":"Notation: The manuscript should clarify whether the translationese classifier is trained separately per language pair or pooled, as this affects interpretation of cross-language consistency.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive feedback on our manuscript. The comments highlight important considerations for the validity of our fluency proxy and the transparency of our statistical controls. We respond to each major comment below, indicating planned revisions where appropriate.","responses":[{"response":"We acknowledge that the POS n-gram classifier primarily captures syntactic patterns characteristic of translationese rather than the full spectrum of literary fluency, including lexical choice and stylistic naturalness. This syntactic focus is consistent with established translationese detection methods in the literature, where such features serve as reliable indicators of non-original-like text. To address the concern, we will revise the Methods and Discussion sections to explicitly discuss the scope and limitations of this proxy measure. We will also add an ablation analysis comparing the classifier against length-controlled baselines and simpler syntactic features to demonstrate that its predictions capture additional signal. While we lack human fluency ratings for the full 130k-paragraph corpus and cannot collect them within the scope of this revision, we will note this as a valuable avenue for future validation studies.","revision_made":"partial","referee_comment":"[Methods] Methods section (classifier description): The claim that the paragraph-level POS n-gram translationese classifier validly measures target-language fluency (original-likeness) is load-bearing for the negative-correlation result. POS sequences primarily encode syntactic distributions; literary fluency also depends on lexical choice, collocations, and stylistic naturalness. Without validation against human fluency ratings on literary paragraphs or an ablation showing that the classifier is not reducible to length or syntax alone, the observed negative correlation with COMET-KIWI risks being an artifact of shared sensitivity to syntactic or length-related features."},{"response":"We agree that additional details on the length-control procedure will improve transparency and allow readers to assess potential residual confounds. In the revised manuscript, we will report the exact linear regression specification (faithfulness regressed on fluency score, paragraph length, and relevant covariates), include variance inflation factors to check for multicollinearity, and provide residual diagnostics (e.g., summary statistics and representative plots) to confirm that length effects have been adequately addressed. These additions will help substantiate that the observed negative correlations are not artifacts of incomplete length control.","revision_made":"yes","referee_comment":"[Results] Results (correlation tables/figures): The length control is described at a high level, but the manuscript does not report the exact regression specification, variance inflation factors, or residual diagnostics. If residual length effects or genre-specific syntactic patterns remain, they could induce the reported negative correlation independently of any genuine fluency-faithfulness tradeoff."}],"tokens_in":1366,"tokens_out":554,"duration_ms":45419,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is a negative correlation between a POS n-gram translationese classifier used as a fluency proxy and COMET-KIWI faithfulness scores, after they control for paragraph length. The pattern shows up clearly in human translations and Google Translate but comes out weaker and often non-significant for TranslateGemma outputs. They work with 130k paragraphs from 106 novels across 16 languages, which is a decent scale for literary material.","headline":"The paper reports a negative correlation between POS-n-gram fluency and COMET-KIWI faithfulness after length control in literary translations, with the pattern stronger for human and Google outputs than for TranslateGemma.","tokens_in":2312,"tokens_out":173,"would_cite":false,"duration_ms":38310,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"Fluency is measured as original-likeness with a translationese classifier trained on paragraph part-of-speech n-grams, and faithfulness with the automatic translation evaluation metric COMET-KIWI."},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/ArithmeticFromLogic.lean","rs_theorem":"embed_strictMono_of_one_lt","paper_passage":"We control for paragraph length and find a consistent negative correlation between fluency and faithfulness."}],"headline":"Linguistic translation metrics (POS n-gram translationese + COMET-KIWI) unrelated to RS recognition-cost or distinction-forcing machinery","alignment":"orthogonal","rationale":"Paper operationalizes fluency via POS n-gram logistic regression on literary paragraphs and faithfulness via COMET-KIWI neural scores, then computes length-controlled partial Spearman correlations. No reference to J-cost, phi-ladders, 8-tick periodicity, or any RS theorem; domain is empirical NLP evaluation of human/LLM literary output. RS modules (AbsoluteFloorClosure, Cost/FunctionalEquation, AlexanderDuality, etc.) address distinguishability-to-spacetime forcing and reciprocal-cost uniqueness; none intersect translationese classification or adequacy metrics.","tokens_in":48095,"confidence":"high","tokens_out":328,"duration_ms":11152,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Literary translations that sound more natural in the target language tend to preserve less of the source meaning.","keywords":["literary translation","fluency","faithfulness","machine translation","translationese","COMET-KIWI","novel translation","automatic evaluation"],"falsifier":"A replication on a comparable set of literary novel paragraphs that finds no negative correlation, or a positive correlation, between the fluency classifier scores and COMET-KIWI scores after length control.","tokens_in":2540,"feed_emoji":"📖","tokens_out":650,"duration_ms":26200,"temperature":0.7,"pith_summary":"The paper tests whether fluency in literary translation comes at the cost of faithfulness to the original text. Using over 130,000 paragraphs from 106 novels across 16 languages, it compares human translations with outputs from Google Translate and TranslateGemma. Fluency is scored by how closely a paragraph matches typical target-language patterns via part-of-speech n-grams, while faithfulness is scored by the COMET-KIWI metric. After accounting for paragraph length, the analysis finds a negative correlation between the two measures for human work and Google Translate, with a weaker pattern for TranslateGemma.","feed_headline":"Natural literary translations often drift from the original meaning","feed_subtitle":"Analysis of 130,000 novel paragraphs shows fluent target-language versions lose more source semantics, even in human work.","key_machinery":"A paragraph-level part-of-speech n-gram classifier that measures original-likeness as a proxy for fluency, paired with the COMET-KIWI metric for semantic faithfulness, applied to a controlled corpus of novel paragraphs.","core_discovery":"The central claim is that fluency and faithfulness trade off against each other in literary novel translation. When paragraphs are made to resemble original writing in the target language, they tend to diverge more from the semantic content of the source, a pattern that holds after controlling for length and appears in both human translations and Google Translate but is reduced in TranslateGemma.","pith_inferences":["The same negative relationship might appear in other genres if measured at paragraph scale, though the paper does not test this.","Translation tools could be designed to let users explicitly choose points along the fluency-faithfulness curve rather than optimizing for one at the expense of the other.","Paragraph-level analysis may reveal different dynamics than sentence-level evaluation in future studies of machine translation."],"forward_implications":["Automatic evaluation of literary translations should account for segment length because it influences the observed fluency-faithfulness relationship.","Human translators and established machine systems exhibit similar tradeoffs when rendering novels.","Newer LLM-based translators may reduce the strength of the tradeoff compared with earlier systems.","The observed pattern suggests that improving fluency metrics alone may not improve overall quality for literary text."],"fun_headline_variants":["Fluency trades off against faithfulness in literary translations","Literary translation fluency reduces source faithfulness","Negative fluency-faithfulness link in human and machine translations","Length controls confirm fluency harms faithfulness in novels"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The part-of-speech n-gram classifier truly captures target-language fluency and COMET-KIWI truly captures semantic faithfulness for paragraphs drawn from novels.","fun_headline_variants_meta":{"raw":{"variants":["Fluency trades off against faithfulness in literary translations","Literary translation fluency reduces source faithfulness","Negative fluency-faithfulness link in human and machine translations","Length controls confirm fluency harms faithfulness in novels"]},"model":"grok-4.3","cost_usd":0.006763,"raw_usage":{"total_tokens":3025,"prompt_tokens":586,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":67628000,"prompt_tokens_details":{"text_tokens":586,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2383,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":586,"tokens_out":56,"duration_ms":38587,"temperature":1.0,"reasoning_tokens":2383,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-19T16:07:39.750494+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication on a comparable set of literary novel paragraphs that finds no negative correlation, or a positive correlation, between the fluency classifier scores and COMET-KIWI scores after length control.","supporting_citations":[],"review_version":1}