{"id":"1a064c51-cc90-4613-be3a-c07c1c04d7bf","arxiv_id":"2504.16601","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 34-translation pilot, Google Translate, Bing, and DeepL generally beat GPT-4o, LLAMA-3.1, and GEMMA-2 on BLEU, CHR-F, and METEOR for medical consultation summaries, with LLMs strongest for Vietnamese and Chinese simple texts.","lead":"Three large language models and three traditional machine translation tools were compared on translating two English medical consultation summaries into Arabic, Chinese, and Vietnamese. Traditional tools scored higher on automatic metrics, especially for complex clinician texts, while LLMs were competitive in Vietnamese and Chinese simple texts; the authors conclude current metrics miss clinical relevance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reference translations may be MT-influenced (Methods admits this), so BLEU/chrF/METEOR scores could favor traditional MT purely because the reference was seeded by such tools; a second, MT-free reference is needed.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the reference translations may have been influenced by machine-generated drafts, which would bias metric scores toward tools resembling those drafts. The paper itself acknowledges this possibility in the Methods section, so the concern is grounded in the manuscript rather than imported from outside. The central claim about traditional MT outperforming LLMs is plausible from the reported numbers, but it is not secured against this confound; only a small fraction of cells are discussed in the text, Figure 1 is not available in the arXiv version, and no statistical uncertainty or error analysis is reported. Given that this is explicitly a pilot study, the conditional verdict is appropriate: the empirical comparison can be reported with the caveat, but the headline conclusion depends on a validation step that has not been performed. I therefore recommend no change to the reader's CONDITIONAL verdict, with the concrete MT-free-reference re-evaluation as the decisive check.","tokens_in":9628,"tokens_out":4200,"duration_ms":41186,"concrete_test":"Commission an independent, certified professional translation of the same two English source summaries under explicit instructions that no machine translation tool or MT-assisted drafting be used. Recompute BLEU, chrF, and METEOR for all 34 generated translations against this MT-free reference, and also report scores against the original reference. Then compare system rankings and the traditional-MT-vs-LLM gap between the two references. If the traditional-MT advantage disappears, narrows materially, or reverses (e.g., if Google Translate's Arabic-complex BLEU drops below GPT-4o's), the reference-contamination concern lands and the abstract needs to be weakened. If rankings are stable across references, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that traditional MT tools generally outperformed LLMs on surface-level metrics, rests entirely on BLEU/chrF/METEOR similarities to one professional reference translation per language and summary. The Methods section 'Reference Translations' explicitly concedes that translation-industry practice often involves using machine-generated drafts as a starting point, that ISO 17100 permits this, and that 'some degree of influence from machine-generated drafts might still persist in the reference translations.' If the certified translators seeded their work with Google Translate, Bing, or DeepL and then post-edited it, the reference is not independent ground truth: it is partly an MT output. Surface metrics such as BLEU and chrF reward lexical and character overlap, so the systems that produced the seed (or outputs resembling it) receive inflated scores, while LLMs that produce fluent paraphrases are penalized. The paper gives no information on which MT engine(s), if any, were used, so the direction and magnitude of this bias cannot be quantified from the text. Because no human clinical evaluation, no error analysis, and no alternative-reference comparison is provided, this single admitted confound is the most load-bearing threat to the abstract's empirical conclusion. This is an evaluation-design concern, not a claim of misconduct, and the paper is honest in flagging it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This pilot study compares three large language models (GPT-4o, GEMMA-2, LLAMA-3.1) with three traditional machine translation tools (Google Translate, Microsoft Bing Translator, DeepL) for translating two English medical consultation summaries into Arabic, Chinese, and Vietnamese. One summary is a simple, patient-facing document and the other is a complex, clinician-oriented letter with medical jargon. Translation quality is measured with BLEU, chrF, and METEOR against professional third-party reference translations. The authors report that traditional MT tools generally scored higher on surface-level metrics, especially for the complex summaries, while some LLMs achieved comparable or higher METEOR scores in Vietnamese and Chinese. They also report language-specific trends, such as Chinese showing a larger score drop from simple to complex summaries and Arabic showing relative improvement on the complex summary. The paper concludes that current automatic metrics are insufficient for clinical translation quality, that LLMs remain inconsistent, and that human oversight is necessary.","tokens_in":9823,"tokens_out":6528,"duration_ms":65841,"significance":"If the empirical pattern holds, the study offers a timely, domain-specific caution that general-purpose LLMs may currently lag behind established MT services on lexical fidelity for medical documents in low- and medium-resource languages, while also indicating that LLMs can be competitive on metrics that tolerate paraphrase. The strengths of the paper are its transparent design, the use of two expert-authored fictitious consultation summaries to avoid privacy issues, the engagement of a certified translation service for reference translations, and the candid acknowledgment of limitations in both the reference construction and the automatic metrics. However, because each condition is a single output scored against a single presumably independent reference, and because the reference may have been influenced by machine translation drafts, the quantitative conclusions should be treated as preliminary and requiring stronger validation before informing clinical practice.","major_comments":[{"comment":"The paper explicitly admits that 'some degree of influence from machine-generated drafts might still persist in the reference translations' (Methods, Reference Translations). This is a load-bearing confound for the abstract's central claim that traditional MT tools generally outperformed LLMs on surface-level metrics: BLEU and chrF reward lexical and character overlap, so a reference that was seeded or post-edited from Google Translate, Bing, or DeepL will inflate the scores of systems whose outputs resemble those drafts and penalize fluent LLM paraphrases. The manuscript does not state which MT engines, if any, the translators used, nor how much post-editing occurred, so the direction and magnitude of the bias cannot be assessed. The authors should either obtain an MT-free reference translation or a second reference, or explicitly reframe every metric-based conclusion as relative to the specific professional references used, rather than as a general statement about translation quality.","section":"Methods, Reference Translations"},{"comment":"The evaluation consists of 34 translations with exactly one output per condition and one reference translation, and the paper reports no variance estimates, confidence intervals, or significance tests. For example, in the Vietnamese simple summary Google Translate's BLEU of 0.7719 is compared with LLAMA's 0.7517, a difference of 0.02 on a single document; without bootstrap over sentences, multiple independent references, or repeated generation runs, such differences cannot be distinguished from noise. The conclusion that traditional MT tools 'generally outperformed' LLMs is therefore not statistically supported. The authors should add uncertainty quantification, or explicitly restrict all conclusions to descriptive comparisons of the particular outputs shown in Figure 1.","section":"Methods, Statistical Analysis; Table 1"},{"comment":"The comparison between the simple and complex summaries is confounded with document identity: the two English source texts differ not only in complexity but also in content, length, and clinical scenario. Therefore statements such as 'Arabic translations improved with complexity due to the language's morphology' and 'Chinese showed the most performance decline with increased complexity' are not supported as causal claims about complexity. To support such conclusions, the authors would need multiple paired documents varying complexity while holding content constant, or a substantially larger sample of source documents. As it stands, these are descriptive differences between two specific texts, and the causal language in the Abstract and Results should be softened accordingly.","section":"Results, Simple vs. Complex Summary Translation Performance"},{"comment":"Table 2 and the Statistical Analysis section describe METEOR as measuring 'similarity at the semantic level' and offering 'a balance between surface and semantic matching.' In fact, METEOR aligns surface word forms with optional exact, stem, and synonym matching; it does not compute meaning-based semantic similarity. Consequently, the Results claim that LLAMA and GEMMA achieve 'competitive or superior METEOR scores' in Vietnamese and Chinese should not be interpreted as evidence of superior semantic preservation. The metric description and any language about semantic accuracy in the Discussion need to be revised to reflect METEOR's actual surface-oriented behavior.","section":"Methods, Statistical Analysis; Table 2"},{"comment":"The abstract draws conclusions about 'medical consultation summaries' generally, but the empirical base is exactly one simple and one complex fictitious summary per language. This single pair of source documents makes it impossible to separate system-level performance from document-specific effects, especially given the large variations in score across the two documents. While a pilot study may reasonably use a small sample, the conclusions should be explicitly limited to the two sample texts, and the Discussion should avoid generalizing the relative ranking of systems to other consultation summaries without further data.","section":"Methods, Original Summaries in English; Abstract"}],"minor_comments":[{"comment":"The Data Availability Statement says the data are 'available online via the link,' but no URL appears anywhere in the manuscript; this prevents independent verification and replication.","section":"Data Availability Statement"},{"comment":"The paper uses both 'CHR-F' and 'chrF' inconsistently; the original metric name is chrF (Popović, 2015), and the notation should be unified.","section":"Throughout"},{"comment":"The reference list contains two entries labeled Wang et al. 2023a and Wang et al. 2023b with the same title 'Document-level machine translation with large language models'; this appears to be a duplicated reference.","section":"References"},{"comment":"The phrase 'both patient, friendly and clinician, focused texts' contains a punctuation and spelling issue; it should be 'patient-friendly and clinician-focused texts.'","section":"Abstract"},{"comment":"Figure 1 is described as presenting the comparative evaluation, but no full table of the 34 metric scores is provided; adding a numeric table would substantially improve the reproducibility and verifiability of the reported comparisons.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a small, honest pilot study, and the authors are to be credited for acknowledging the limitations of their reference translations and evaluation metrics. The main unresolved risk is that the reference translations may have been produced with the assistance of the very MT tools being evaluated, which would directly bias the central comparison; the authors should be asked to disclose the tools used by the translation vendors and, if possible, to provide an MT-free reference or a contamination analysis. In addition, the cross-document complexity comparisons and the lack of uncertainty quantification need to be addressed before the findings can be considered reliable enough for a clinical audience."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a small, honest pilot, and the headline result—traditional MT ahead of LLMs on BLEU/chrF, less so on METEOR—is plausible but not robust. The sample is two documents, three languages, one reference per cell, and the reference translations may have been seeded by MT. That last point is admitted in the Methods, which is to the authors' credit, but it undercuts the abstract's confident phrasing.\n\nWhat's new: I don't know of another comparison of GPT-4o, LLAMA-3.1, GEMMA-2, Google Translate, Bing, and DeepL on palliative care summaries into Arabic, Chinese, and Vietnamese. The Arabic-complexity pattern (scores rise with complexity) is fresh and worth a follow-up. The paper is clearly written, the translation protocol is described in enough detail to replicate, and the limitations they list are genuine rather than perfunctory.\n\nWhere it's soft: n=2 per language, no multiple references, no error bars or significance tests, no human clinical evaluation. The claim that 'current metrics fail to capture clinical relevance' is reasonable but it's an opinion, not a finding from this data. Similarly, the discussion of LLM limitations (terminology mapping, prompt sensitivity) is general and not tied to specific outputs. The missing figure and data link in the arXiv version make independent checking harder. The reference contamination is the sharpest issue: if the certified translators started from Google/Bing/DeepL drafts and post-edited, the surface metrics are biased in favor of those engines. The authors acknowledge this but do not quantify it. A second MT-free reference, or human scoring on error severity, would fix it.\n\nBottom line: this deserves peer review as a pilot. The appropriate outcome would be a revise-and-resubmit that adds hedging to the abstract, either adds the missing figure/data link or explains their absence, and frames the system ranking as provisional. I wouldn't cite it in my own work, but I'd read a follow-up with a real corpus.","headline":"Honest, tiny pilot: traditional MT edges out LLMs on surface metrics, but n=2 docs and possibly MT-seeded references keep the headline provisional.","tokens_in":10390,"tokens_out":2778,"would_cite":false,"duration_ms":26504,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Traditional machine-translation services generally outperform large language models on automated fidelity metrics for medical consultation summaries in Arabic, Chinese, and Vietnamese, especially for complex clinician-facing text, a pilot…","keywords":["machine translation","large language models","medical translation","consultation summary","responsible AI","BLEU","CHR-F","METEOR"],"falsifier":"Translate a larger corpus of consultation summaries—say 50 simple and 50 complex per language—with the same six systems, score them against independently produced human references verified to contain no machine-translation influence, and check whether LLMs match or beat Google Translate, Bing, and DeepL on BLEU, CHR-F, and METEOR for complex texts; if they do, the paper's claim of a general traditional-MT advantage is overturned.","tokens_in":9398,"feed_emoji":"🩺","tokens_out":12830,"duration_ms":107051,"temperature":0.7,"pith_summary":"This pilot study compares three general-purpose large language models—GPT-4o, LLAMA-3.1, and GEMMA-2—with three established machine-translation services—Google Translate, Microsoft Bing Translator, and DeepL—on translating two simulated medical consultation summaries from English into Arabic, simplified Chinese, and Vietnamese: a simple, patient-facing summary and a complex, clinician-oriented letter. Its central finding is that the traditional MT services generally score higher on surface-fidelity metrics—word-level and character-level overlap with professional references—especially for the complex summary, while some LLMs match or beat them on METEOR, a metric that aligns on meaning and tolerates paraphrasing, for Vietnamese and Chinese on the simpler patient text. Arabic runs the opposite way: scores rise with text complexity, which the authors attribute to richer morphological context aiding disambiguation. The paper argues that automated metrics do not capture clinical adequacy and that human oversight remains necessary. If correct, the study indicates that freely available general-purpose LLMs are not yet safe substitutes for domain-vetted or human-reviewed translation in healthcare.","feed_headline":"Translation tools still outscore LLMs on medical summaries","feed_subtitle":"For Arabic, Chinese and Vietnamese medical text, Google, Bing and DeepL edge out GPT-4o, LLAMA and GEMMA on fidelity.","key_machinery":"The central machinery is the reference-anchored metric triad: BLEU counts word-level n-gram overlap, CHR-F counts character-level n-gram overlap, and METEOR aligns words with synonym and stemming tolerance to approximate semantic similarity. All three are computed against professional third-party reference translations of two deliberately contrasted English summaries—a simple lay-language patient summary and a complex jargon-dense clinician letter—across three languages. The simple-versus-complex contrast is the controlled variable that exposes language-specific behavior, such as Arabic benefiting from additional context and Chinese degrading sharply on technical content.","core_discovery":"The paper asserts that when medical consultation summaries are translated from English into Arabic, simplified Chinese, and Vietnamese, traditional MT tools—Google Translate, Microsoft Bing Translator, and DeepL—generally outperform the tested LLMs—GPT-4o, LLAMA-3.1, and GEMMA-2—on surface-level metrics (BLEU and CHR-F), with the advantage most pronounced for the complex clinician-oriented summary. Against that pattern, LLAMA-3.1 and GEMMA-2 achieve METEOR scores in Vietnamese and Chinese comparable to or above the traditional tools on the simpler patient-facing summary, and Arabic translation quality improves with complexity across most systems because longer, context-rich sentences help resolve the language's morphological ambiguities. Chinese shows the steepest drop from simple to complex text, which the authors tie to syntactic and terminological challenges. The paper presents these results as an indicative pilot comparison, not a definitive quality ranking, and stresses that the metrics cannot measure clinical safety.","pith_inferences":["Editorial inference: the paper acknowledges that the professional reference translators may have used machine-generated drafts, so the metric scores could systematically favor outputs resembling those drafts; scoring against independently produced human references would test whether the traditional-MT advantage persists.","Editorial inference: because each language is represented by only one simple and one complex summary, the language-level patterns are provisional; a larger corpus of consultation summaries would show whether Vietnamese resilience and the Arabic complexity gain are stable.","Editorial inference: a clinician-oriented error-severity evaluation—counting mistranslated medication names, dosages, and instructions—would likely rank the systems differently from BLEU, CHR-F, and METEOR and give a more actionable safety signal for deployment.","Editorial inference: the Vietnamese METEOR results suggest fine-tuning LLMs on biomedical Vietnamese corpora is a plausible next step, but the paper does not test whether such fine-tuning would close the surface-fidelity gap with traditional MT."],"forward_implications":["Healthcare workers using a free default tool on complex clinician-oriented text would currently get closer surface-level fidelity from Google Translate or Bing than from the tested general-purpose LLMs.","For patient-facing Vietnamese and Chinese, an LLM such as LLAMA-3.1 or GEMMA-2 can produce translations whose semantic similarity to a professional reference rivals or exceeds traditional MT, making them plausible draft tools for simpler text.","Automatic metric scores cannot certify a medical translation as safe: a mistranslated drug name and a missing filler word can receive the same penalty, so clinical review remains necessary before use.","Arabic-language medical translation appears to improve with longer, more context-rich input, a pattern that runs opposite to the Chinese and Vietnamese results."],"supporting_citations":[{"why":"Supplies BLEU, the word-overlap metric used to score every translation.","marker":"Papineni et al. (2002)"},{"why":"Supplies CHR-F, the character-level metric used to measure morphological fidelity.","marker":"Popović (2015)"},{"why":"Supplies METEOR, the semantic-alignment metric used to approximate meaning preservation.","marker":"Banerjee and Lavie (2005)"},{"why":"Establishes the medical-domain translation challenges and safety stakes motivating the comparison.","marker":"Costa-Jussà et al. (2012)"},{"why":"Provides the census data that identify Arabic, Chinese, and Vietnamese as the most spoken non-English languages in Australia.","marker":"Australian Bureau of Statistics (2022)"},{"why":"Makes the document-level claim that LLMs can outperform traditional MT, the expectation the study tests against.","marker":"Wang et al. (2023b)"},{"why":"Specifies Google Translate as one of the traditional MT systems whose outputs are scored.","marker":"Google (2024)"},{"why":"Specifies Microsoft Bing Translator as a second traditional MT system in the comparison.","marker":"Microsoft Corporation (2024)"},{"why":"Specifies DeepL as the third traditional MT system and explains why Vietnamese translations were unavailable from it.","marker":"DeepL (2024)"}],"fun_headline_variants":["MT tools outperform LLMs on medical translation","LLMs lose to traditional MT on medical summaries","Medical translation: MT beats LLMs in Arabic, Chinese, Vietnamese","Google, Bing, DeepL top GPT-4o, LLAMA, GEMMA on medical text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the professional reference translations are an unbiased ground truth, even though the paper concedes translators may have used machine-generated drafts as a starting point—if those references were influenced by MT output, the scores are biased toward systems resembling those drafts.","fun_headline_variants_meta":{"raw":{"variants":["MT tools outperform LLMs on medical translation","LLMs lose to traditional MT on medical summaries","Medical translation: MT beats LLMs in Arabic, Chinese, Vietnamese","Google, Bing, DeepL top GPT-4o, LLAMA, GEMMA on medical text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1255,"prompt_tokens":860,"completion_tokens":395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":322}},"tokens_in":476,"tokens_out":395,"duration_ms":4164,"temperature":1.0,"reasoning_tokens":322,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:58:44.300045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Translate a larger corpus of consultation summaries—say 50 simple and 50 complex per language—with the same six systems, score them against independently produced human references verified to contain no machine-translation influence, and check whether LLMs match or beat Google Translate, Bing, and DeepL on BLEU, CHR-F, and METEOR for complex texts; if they do, the paper's claim of a general traditional-MT advantage is overturned.","supporting_citations":[],"review_version":1}