{"id":"c39e29b2-f078-478d-8e34-85a0a809be0c","arxiv_id":"2505.01560","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Agentic machine translation consumes 5 to 15 times more tokens than NMT or single-pass LLMs without delivering consistently higher quality in this small benchmark, while o1-preview tops human ratings.","lead":"This paper benchmarks Google Translate, two LLMs, and two AI-agent translation workflows on legal and news texts in three language pairs, measuring translation quality and token costs. The agent workflows use 5 to 15 times more tokens than NMT or single-pass LLMs without clearly better output, while a reasoning LLM leads in human ratings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human evaluation is too underpowered to support the quality-reversal claim: Table 2 means differ by as little as 0.03 on a 4-point scale, with no reported rater count or error bars, so the 'not worth resources' conclusion rests on possibly noisy rankings.","rationale":"The reader identified the human evaluation as the weakest assumption, and I agree. The paper's headline conclusion is a cost-quality trade-off: agents are not worth the resources because the quality gain is absent. The cost side is well measured and reproducible from Table 3, but the quality side has two pillars, both on tiny samples. The automatic metrics use only one legal contract (537 words) and one news article (116 words) per language, so metric differences of a few tenths are within noise. The human evaluation is meant to supply the validity that automatic metrics lack, yet it is described so thinly that no statistical inference is possible. The abstract's 'five of six' and 'edges ahead once' are presented as findings, but with no rater count, no segment count, and no reliability measure, these rankings could easily be reordered by a different expert panel. This matters directly for the central claim: if a properly powered human evaluation found that the iterative agent or sequential agent significantly improves adequacy or fluency, the 'not worth the resources' conclusion for those workflows would weaken. Conversely, if the current findings are just noise, the paper overstates its negative result. The authors deserve credit for acknowledging the i-agent technical limitations and for providing token counts that support the multiplicative cost claim; that part of the paper stands independently. The appropriate remedy is to require the authors to release the raw ratings and provide confidence intervals or significance tests, and to temper the abstract's reversal language until that is done. Hence the conditional verdict should remain unchanged.","tokens_in":13036,"tokens_out":9398,"duration_ms":97564,"concrete_test":"Request the raw segment-level human ratings and the number of raters for Section 3.3/Table 2. For each language and dimension, bootstrap the per-system means (10,000 resamples) and compute 95% percentile confidence intervals for all pairwise differences, e.g., o1-preview vs GT in Spanish adequacy (3.92 vs 3.81). If any interval separating a claimed best from a claimed runner-up contains zero, the human-eval reversal is not statistically supported; additionally compute the minimum number of raters/segments needed to detect the observed effect at α=0.05 with 80% power. If that N is not satisfied by the actual data, the quality conclusion cannot be distinguished from noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central recommendation—that agentic MT is not yet worth its token cost—depends on showing that the extra resources buy no consistent quality gain. The automatic metrics (Table 1) are computed on two short documents (537 and 116 words) per language, so their small score gaps are not statistically robust. The more consequential evidence, the human evaluation (Section 3.3, Table 2), is the only measure intended to capture quality invisible to surface metrics, and it is used in the abstract to claim a partial reversal: o1-preview is best in five of six comparisons and the iterative agent edges ahead once. However, Table 2 reports no number of raters, no number of rated segments, no inter-rater reliability, and no significance tests. The winning margins are frequently tiny—e.g., Spanish adequacy 3.92 vs 3.81 (o1 vs GT), Catalan fluency 3.69 vs 3.61 (o1 vs i-agent), Spanish fluency 3.72 vs 3.64 (i-agent vs GPT-4o). On a 4-point scale with likely few observations, these differences are well within sampling noise. If the human rankings are noise, the 'quality reversal' narrative collapses, and the conclusion that agents provide no quality benefit is unsupported; the paper would then only have a cost measurement, which alone cannot justify the title's claim. The cost ratios themselves (Section 4.3) are more robust, and the authors' explicit admission that the iterative agent had technical limitations is creditable, but that admission further weakens the i-agent quality data rather than rescuing it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks five machine translation paradigms—Google Translate (NMT), GPT-4o (general-purpose LLM), o1-preview (reasoning-enhanced LLM), a sequential three-stage agent (s-agent), and an iterative refinement agent (i-agent)—on one legal contract (537 words) and one news article (116 words) translated from English into Spanish, Catalan, and Turkish. Quality is evaluated with COMET, BLEU, chrF2, and TER, plus expert human ratings of adequacy and fluency on a 4-point scale; efficiency is measured as input-plus-output token counts mapped to April 2025 prices. The central findings are that Google Translate leads most automatic metric comparisons, that human evaluation reverses part of that narrative by favoring o1-preview in five of six comparisons and the i-agent once, and that the s-agent consumes roughly five times and the i-agent roughly fifteen times the tokens of single-pass systems, with correspondingly higher monetary cost. The paper argues that agentic MT is not yet worth its resource premium and recommends cost-aware, multidimensional evaluation protocols.","tokens_in":13354,"tokens_out":2509,"duration_ms":27527,"significance":"If the empirical claims held with adequate statistical support, the paper would be a useful reality check on the current hype around LLM agents for machine translation: the token-cost ratios for sequential and iterative agent workflows are large, robust, and practically relevant for deployment decisions, and the observation that automatic metrics and human ratings diverge for reasoning-enhanced LLMs is a timely methodological point. The cost measurement itself, based on token counts, is simple, transparent, and credible, and the authors' explicit admission that their iterative agent suffered technical limitations is a sign of good faith. However, the paper's headline quality conclusions rest on human ratings and automatic scores computed over an extremely small sample, with no reported rater number, inter-rater reliability, or significance testing. Because the quality-reversal narrative and the 'not yet worth the resources' conclusion depend on these underpowered measurements, the central interpretive claim is currently load-bearing on noise-prone evidence.","major_comments":[{"comment":"The human evaluation is too underpowered to support the quality-reversal claim. Table 2 reports no number of raters, no number of rated segments, no inter-rater reliability, and no significance tests, yet the abstract and conclusion use these data to claim that o1-preview is best in five of six comparisons and that the iterative agent edges ahead once. Several winning margins are very small on a 4-point scale—for example, Spanish adequacy 3.92 vs 3.81, Catalan fluency 3.69 vs 3.61, and Spanish fluency 3.72 vs 3.64—and with a small number of expert ratings these differences are well within sampling noise. The authors should either report the evaluation design and demonstrate that the differences are statistically reliable, or soften the conclusion to describe the human ratings as suggestive rather than probative.","section":"§3.3, Table 2"},{"comment":"The automatic evaluation is computed on two short documents (537 and 116 words) per language pair, so the score gaps in Table 1—often one or two COMET/BLEU points—are not demonstrated to be statistically meaningful. The paper's claim that 'GT ranks first in seven of twelve metric-language combinations' is a descriptive ranking of point estimates with no confidence intervals or resampling, and the subsequent interpretive paragraph about GT's 'dominance' overstates what can be inferred from a single legal contract and a single news article. The authors should either provide uncertainty estimates for the automatic scores or clearly frame the results as a pilot study, which would also align with the title's phrase 'initial exploration.'","section":"§3.2, Table 1"},{"comment":"The cost analysis omits o1-preview entirely because token counts were not available, yet o1 is one of the five benchmarked systems and the system that most often wins the human evaluation. The paper's inference that o1's token consumption 'would be higher than that of GPT-4o' is speculation, not measurement, and it prevents any quantitative statement about the cost-quality trade-off for the best-performing system in the human evaluation. Without o1 token data, the 'steep costs' narrative applies only to the two agent workflows, and the title's implied resource-vs-quality verdict for reasoning-enhanced LLMs is unsupported.","section":"§4.3, Table 3"},{"comment":"The Discussion concedes that the iterative agent had technical limitations and that 'the results could have been substantially better with a more adequate development.' This admission weakens the comparative quality findings for the i-agent and should be disclosed earlier, in the methodology section, so that readers can interpret Table 1 and Table 2 accordingly. As written, the paper evaluates a possibly broken implementation and then draws conclusions about iterative agent workflows generally.","section":"§5, Discussion"}],"minor_comments":[{"comment":"The news document is described as having a 'type-token ration of 0.654'; 'ration' should be 'ratio.'","section":"§3.2"},{"comment":"The table uses 'o1--preview' with a double hyphen in the Spanish and Turkish rows, while the Catalan row and the text use 'o1-preview'; please standardize.","section":"§4.1, Table 1"},{"comment":"The legal contract is cited as '(ANONYMISED)' without a specific reference or a link to the source text; since the authors emphasize reproducibility, the source should be identified or the text should be shared in an appendix.","section":"§3.2"},{"comment":"The USD price column for Google Translate lists '$20 (chars.)', which is not directly comparable to the per-million-token prices for the API-based systems; the text should clarify whether the Google price is per million characters and explain the basis of the $0.038 cost estimate for the sample documents.","section":"§4.3, Table 3"},{"comment":"In the paragraph on Turkish, the statement that 'a closer examination of the s-agent and i-agent results' revealed grammatical issues suggests a qualitative error analysis, but no error typology, example outputs, or detail is provided; either include the analysis or mark it as future work.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a small-scale empirical study with a clear and relevant cost-efficiency message, but the quality evaluation is not yet statistically grounded. The most consequential fix is to either supply the missing human-evaluation metadata and significance tests or explicitly demote the quality claims to pilot-level observations. I would also recommend that the authors either obtain o1 token counts or drop the cost-quality comparison for o1 from the conclusions. The automatic-metric section and the cost table are useful and could survive revision; the current evidence does not justify the title's resource-versus-quality verdict without these changes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this paper gives you a clean, credible token-cost comparison of multi-agent MT against baselines, and that part is worth reading. The quality conclusions, including the title's \"not worth the resources,\" rest on a human evaluation that is too thin to carry the weight.\n\nWhat's new: the side-by-side measurement of sequential and iterative agent workflows against Google Translate, GPT-4o, and o1-preview, holding the underlying model constant and using the same prompts. The cost ratios are stark: ~5x and ~15x token consumption. That's a useful data point for anyone advising translation buyers.\n\nThe paper does some things right. The setup is transparent about the two agent architectures, the token accounting is simple to follow, and the authors explicitly flag that their iterative agent had technical problems—that honesty is creditable. The broader argument for cost-aware evaluation is reasonable.\n\nThe soft spots are where the stress-test note lands. The human evaluation (Table 2) reports no number of raters, no inter-rater reliability, no significance tests, and the winning margins are often 0.03–0.11 on a 4-point scale. On two short documents per language pair, those differences are sampling noise. The automatic metrics (Table 1) are also computed on a 537-word contract and a 116-word news piece, so the rankings there are descriptive at best. The paper's abstract and conclusion treat o1-preview's human-rated edge as evidence that reasoning layers \"capture semantic nuance undervalued by surface metrics.\" That may be true, but this study doesn't have the statistical power to show it.\n\nA second, smaller issue: o1-preview token counts are absent (no API access), so the cost comparison excludes one of the five systems. The authors infer it would fall between GPT-4o and the agents, which is plausible but not measured.\n\nNet: the cost finding is real and useful; the quality-reversal narrative is not supported by the evidence as presented. I would send this to peer review with a directive to strengthen the human evaluation (or drop the reversal claim), report rater details, and add error bars or significance tests. The cost measurement alone gives the paper enough value to merit referee time.\n\nRecommendation: accept for review with expected heavy revision.","headline":"Credible token-cost data on multi-agent MT, but the quality-reversal conclusion rests on a human evaluation too underpowered to support it.","tokens_in":13884,"tokens_out":2208,"would_cite":false,"duration_ms":21764,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multi-agent AI translation workflows consume about five to fifteen times as many tokens as single-pass systems while failing to beat a standard NMT baseline on automatic metrics.","keywords":["machine translation","AI agents","multi-agent systems","large language models","cost-efficiency","human evaluation","automatic metrics","reasoning-enhanced LLM"],"falsifier":"Re-run the human evaluation with a larger set of texts and multiple raters, reporting per-rater scores and inter-rater agreement; if the near-tied differences between o1-preview and the other systems do not replicate, the claimed human-quality advantage is not stable.","tokens_in":12845,"feed_emoji":"🌐","tokens_out":6840,"duration_ms":62026,"temperature":0.7,"pith_summary":"This paper puts the hype around AI-agent translation to an empirical test by benchmarking five ways of translating the same English legal and news texts into Spanish, Catalan, and Turkish. The results show that a mature NMT system (Google Translate) still wins seven of twelve automatic metric-and-language comparisons, while a reasoning-enhanced LLM (o1-preview) wins five of six human evaluation comparisons. Both multi-agent workflows trail on automatic metrics and consume about five times (sequential) and fifteen times (iterative) as many tokens as single-pass systems. The paper concludes that agentic translation does not yet justify its resource premium, and argues for evaluation protocols that weigh quality, cost, and human judgment together.","feed_headline":"AI translation agents use 5-15x tokens without beating NMT","feed_subtitle":"Benchmarking five systems in three languages, the paper shows agent workflows pay a big token premium without metric gains.","key_machinery":"The machinery is a controlled comparison of five translation pipelines on identical texts and prompts: a production NMT system, a general-purpose LLM, a reasoning-enhanced LLM, and two multi-agent workflows built on the same underlying model. The sequential agent runs a translator, a reviewer, and an editor in fixed order; the iterative agent runs the same roles with up to three refinement cycles. Holding the model and prompts constant isolates architecture as the cause of any quality difference, while the cost measure is total input-plus-output token count mapped to April 2025 prices.","core_discovery":"The central claim is that, in the tested conditions, adding agentic orchestration to an LLM buys little measurable quality at a steep resource price. Google Translate ranks first in seven of twelve automatic metric-language combinations, and o1-preview ties or places second in most of the rest, while the sequential and iterative agents trail; no agent workflow wins an automatic comparison. Human expert ratings reverse part of the picture, with o1-preview rated most adequate and fluent in five of six comparison dimensions and the iterative agent once, suggesting reasoning layers capture nuance that surface metrics miss. Yet the iterative agent consumes roughly fifteen times, and the sequential agent about five times, the tokens used by Google Translate or a single-pass LLM, which the authors interpret as evidence that agentic MT is promising but not yet cost-effective.","pith_inferences":["If token prices continue to fall, the resource premium that currently disqualifies agentic MT may shrink; whether the quality advantage remains is a separate empirical question that this study's design does not answer.","A testable extension would introduce a targeted metric or evaluation rubric for pragmatic adequacy and see whether agent outputs' human-rated edge shows up automatically; the paper's divergence result predicts it would.","The large Turkish gap between the two agent architectures suggests that coordination strategy and morphologically rich targets interact; exploring role prompts specialized by language typology could be a direct next experiment."],"forward_implications":["Under current models and prices, a cost-conscious translation pipeline would still choose the NMT baseline or a single-pass LLM over agentic workflows for general-purpose legal and news text.","Reasoning-enhanced LLMs offer a larger human-perceived quality gain per extra token than multi-agent coordination, because o1-preview reaches top human ratings without iterative loops.","Automatic metrics can mis-rank systems that differ on semantic adequacy; the paper's human results imply metric-only benchmarking under-reports reasoning-driven and agentic quality.","Token budgets of 10,000 to 39,000 for translating short documents make iterative agent pipelines difficult to scale without efficiency measures."],"supporting_citations":[{"why":"Supplies the translation-agent scaffold on which the sequential three-stage workflow is built.","marker":"Ng [2024] 2025"},{"why":"Provides the evaluator-optimizer architecture used for the iterative agent workflow.","marker":"Briva-Iglesias 2025"},{"why":"Underpins the choice and interpretation of the automatic metrics COMET, BLEU, chrF2, and TER.","marker":"Kocmi et al. 2021"},{"why":"Context for treating reasoning-enhanced LLMs as a distinct MT paradigm and for the o1-preview comparison.","marker":"Liu et al. 2025"},{"why":"Grounds the use of human expert evaluation to capture dimensions that automatic metrics miss.","marker":"Freitag et al. 2021"},{"why":"Establishes the baseline expectation that LLMs can approach NMT quality, motivating the head-to-head design.","marker":"Hendy et al. 2023"}],"fun_headline_variants":["Agentic MT: 5-15x tokens, no metric win","Translation agents cost 15x tokens, trail NMT","Human raters favor o1, but token cost explodes","NMT beats agents on metrics, agents cost 15x more","Agent workflows: quality edge small, token cost huge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the small set of expert ratings, with no reported number of raters or inter-rater reliability, can reliably rank systems whose scores differ by as little as 0.03 on a 4-point scale.","fun_headline_variants_meta":{"raw":{"variants":["Agentic MT: 5-15x tokens, no metric win","Translation agents cost 15x tokens, trail NMT","Human raters favor o1, but token cost explodes","NMT beats agents on metrics, agents cost 15x more","Agent workflows: quality edge small, token cost huge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1360,"prompt_tokens":1031,"completion_tokens":329,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":243}},"tokens_in":647,"tokens_out":329,"duration_ms":3795,"temperature":1.0,"reasoning_tokens":243,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:15:32.031986+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the human evaluation with a larger set of texts and multiple raters, reporting per-rater scores and inter-rater agreement; if the near-tied differences between o1-preview and the other systems do not replicate, the claimed human-quality advantage is not stable.","supporting_citations":[],"review_version":1}