{"id":"3fcbbbbb-1b0a-4772-a3aa-0a65cc3e7f7e","arxiv_id":"2605.13368","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Document-level machine translation followed by segment-level LLM refinement provides the strongest and most stable improvements in literary translation quality, mainly enhancing fluency and style rather than adequacy.","lead":"This paper conducts a large-scale study on how iterative self-refinement by LLMs affects document-level literary translations across nine models and seven language pairs. It identifies a reliable pipeline of document-level initial translation followed by segment-level refinement that improves fluency, style, and terminology more consistently than other approaches.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's identification of generalizability as the weakest point is accurate, yet it does not undermine the internal validity of the reported patterns within the tested conditions. The breadth of models, languages, and evaluation methods provides direct empirical support for the headline recipe without detectable gaps in the argument structure.","tokens_in":1696,"tokens_out":246,"duration_ms":25044,"concrete_test":"Recompute the primary quality deltas (fluency/style/terminology vs. adequacy) for the top two granularity combinations on the largest language pair using the same human annotation protocol; if the relative ordering reverses or the effect size drops below statistical significance, the robustness claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The study reports consistent patterns across nine LLMs, seven language pairs, nine granularity combinations, five strategies, and large-scale human evaluation. The central claim—that document-level MT followed by segment-level refinement produces stronger, more stable gains than document-level refinement—rests on these direct comparisons. No internal inconsistency, unsupported derivation, or untested assumption required for the claim to hold is visible in the experimental design or reported outcomes.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a systematic empirical study of iterative self-refinement for document-level literary machine translation with LLMs. It evaluates nine LLMs across seven language pairs using nine granularity combinations and five refinement strategies, supported by large-scale human evaluation. The central claim is that document-level MT followed by segment-level refinement produces strong, stable gains (mainly in fluency, style, and terminology), while document-level refinement yields fewer edits and less reliable improvements; a simple general prompt outperforms error-specific or evaluate-then-refine variants, and refinement aligns outputs to the refiner's distribution rather than targeted error correction.","tokens_in":1759,"tokens_out":344,"duration_ms":34409,"significance":"If the findings hold, the work offers clear practical guidance for LLM refinement pipelines in literary translation and illuminates the mechanisms and limits of current approaches. The broad coverage of models, languages, and strategies, together with consistent patterns from human judgments, provides a solid empirical foundation that can inform both research and deployment of inference-time MT improvements.","major_comments":[],"minor_comments":[{"comment":"Abstract: The abstract states that large-scale human evaluation was performed but does not mention statistical significance tests or inter-annotator agreement; adding one sentence on these points would strengthen the summary of the results.","section":"Abstract"},{"comment":"§5 (or equivalent results section): When reporting the nine granularity combinations, a compact summary table or clearer visual encoding of the exact MT/refinement granularity pairs would make the cross-condition comparisons easier to parse at a glance.","section":"§5"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive review and recommendation for minor revision. We appreciate the recognition that our systematic study provides clear practical guidance for LLM refinement pipelines in literary translation, supported by broad coverage across models, languages, and human judgments.","responses":[{"response":"We thank the referee for this accurate summary of our work. The description aligns closely with our abstract, experimental design, and conclusions. No revisions are required on this point.","revision_made":"no","referee_comment":"The manuscript presents a systematic empirical study of iterative self-refinement for document-level literary machine translation with LLMs. It evaluates nine LLMs across seven language pairs using nine granularity combinations and five refinement strategies, supported by large-scale human evaluation. The central claim is that document-level MT followed by segment-level refinement produces strong, stable gains (mainly in fluency, style, and terminology), while document-level refinement yields fewer edits and less reliable improvements; a simple general prompt outperforms error-specific or evaluate-then-refine variants, and refinement aligns outputs to the refiner's distribution rather than targeted error correction."}],"tokens_in":1214,"tokens_out":244,"duration_ms":40628,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The key point from this paper is that refinement works best when you first produce a full-document translation and then refine it one segment at a time. Full-document refinement tends to make fewer changes and delivers smaller, less consistent gains. The authors also show that the improvements come mostly from fluency, style, and terminology rather than better adequacy, and that the process largely projects the output toward whatever the refiner model would have generated on its own.","headline":"Document-level MT followed by segment-level refinement beats full-document refinement for literary texts, mainly by shifting output toward the refiner model's own distribution rather than targeted error fixing.","tokens_in":2255,"tokens_out":166,"would_cite":true,"duration_ms":18748,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements."},{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy"}],"headline":"Empirical NLP study on LLM MT refinement pipelines; no RS-shaped machinery","alignment":"orthogonal","rationale":"The paper conducts large-scale experiments on translation-refinement granularity (seg/para/doc), prompting strategies, MQM dimension gains (fluency > accuracy), and refiner behavior (ceiling/anchor effects, distribution projection). Its central objects are pipelines, edit ratios, and human MQM scores on WMT24-Literary. RS framework derives J-cost, φ, 8-tick periodicity, D=3, and constants c/ℏ/G from a single distinction (reality_from_one_distinction, AbsoluteFloorClosure, Cost.FunctionalEquation, AlexanderDuality). No overlap exists in objects, forcing relations, or cost functions; the domain (computational linguistics) lies outside RS theorems.","tokens_in":59889,"confidence":"high","tokens_out":346,"duration_ms":10468,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Document-level translation followed by segment-level refinement produces the most reliable gains in literary machine translation.","keywords":["LLM refinement","document-level translation","literary translation","machine translation","self-refinement","granularity effects","translation quality","fluency and style"],"falsifier":"A controlled replication using new LLMs or non-literary documents in which full document-level refinement produces larger and more stable gains than the document-then-segment pipeline would falsify the central recommendation.","tokens_in":2609,"feed_emoji":"🔄","tokens_out":648,"duration_ms":18990,"temperature":0.7,"pith_summary":"The paper tests what iterative LLM self-refinement actually changes when translating full literary documents. Across nine models, seven language pairs, and multiple granularity setups, the strongest and most stable results come from first translating the entire document and then refining it segment by segment. Gains appear mainly in fluency, style, and terminology consistency, while improvements to meaning accuracy remain smaller and less consistent. Refinement also tends to shift the output toward the refiner model's own stylistic distribution rather than repairing specific errors. A plain general prompt works better than prompts that target particular error types.","feed_headline":"Document MT then segment refinement beats full-document fixes","feed_subtitle":"Literary translation experiments show this hybrid pipeline delivers stronger, more consistent gains in fluency and style than document-level","key_machinery":"The document-level MT followed by segment-level refinement pipeline, which carries the argument by separating coarse context handling from fine-grained polishing.","core_discovery":"The central claim is that, for literary translation, an initial document-level machine translation pass followed by segment-level refinement outperforms other granularity combinations and refinement strategies. Document-level refinement produces fewer edits and less reliable quality lifts. Across experiments, refinement improves fluency, style, and terminology more than adequacy, and the process projects the output toward the refiner model's distribution instead of performing targeted error correction. A simple general refinement prompt consistently beats error-specific prompting and evaluate-then-refine schemes.","pith_inferences":["Translation systems may benefit from deliberately separating document-scale context capture from segment-scale polishing in their inference pipelines.","The style-projection finding suggests current refinement has limited power for meaning-level error repair and may need external signals to target adequacy.","The recipe could be tested on other text domains such as technical or conversational material to check whether the granularity preference persists."],"forward_implications":["Refinement gains concentrate on fluency, style, and terminology rather than adequacy.","A single general refinement prompt outperforms error-specific and evaluate-then-refine variants.","The output after refinement moves closer to the refiner model's own distribution.","Document-level initial translation plus segment refinement remains stable across model strengths and language pairs."],"fun_headline_variants":["Document MT then segment refinement yields stable gains","Full document refinement results in fewer edits","Refinement improves fluency and style more than adequacy","Refinement moves output toward refiner model distribution"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The observed superiority of the hybrid granularity recipe and the quality dimension patterns will hold for LLMs, language pairs, and text genres outside the nine models, seven pairs, and literary texts tested.","fun_headline_variants_meta":{"raw":{"variants":["Document MT then segment refinement yields stable gains","Full document refinement results in fewer edits","Refinement improves fluency and style more than adequacy","Refinement moves output toward refiner model distribution"]},"model":"grok-4.3","cost_usd":0.011029,"raw_usage":{"total_tokens":4768,"prompt_tokens":659,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":110290500,"prompt_tokens_details":{"text_tokens":659,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4055,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":659,"tokens_out":54,"duration_ms":26254,"temperature":1.0,"reasoning_tokens":4055,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-14T20:05:46.245306+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled replication using new LLMs or non-literary documents in which full document-level refinement produces larger and more stable gains than the document-then-segment pipeline would falsify the central recommendation.","supporting_citations":[],"review_version":1}